Audio personalization method and system
By establishing an HRTF library and using the user's head characteristics for audio matching, the problem that audio response in the prior art cannot match interactive content is solved, and a personalized 3D audio experience is achieved.
Patent Information
- Application Number
- JP2023519015
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-01
- Filing Date
- 2021-09-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-09-15
AI Technical Summary
The prior art is difficult to achieve audio responses that match interactive content such as video games, and cannot effectively adjust the audio effect according to the user's head characteristics.
By obtaining the head transfer function (HRTF) of multiple reference individuals, an HRTF library is established, and the user's head characteristics are used to personalize audio. Through testing and matching, HRTFs in the library closest to the user's HRTF, and applying it to the user's audio system.
It realizes audio adjustments based on the user's head characteristics, providing a more realistic and personalized 3D audio experience, suitable for all kinds of audio devices and users.
Smart Images

Figure 0007675807000001 
Figure 0007675807000002 
Figure 0007675807000003
Abstract
Description
[Technical field]
[0001] The present invention relates to an audio personalization method and system. [Background technology]
[0002] Consumers of media content, including interactive content such as video games, enjoy a sense of immersion while engaging with the content. With pre-recorded content, there is an implicit understanding that the content is fixed, except for the video and audio. However, for interactive content such as video games, where the content and the perspective of that content typically change based on user input, there is a desire for the audio to respond as well.
[0003] The present invention aims to alleviate or reduce this need. Summary of the Invention
[0004] Various aspects and features of the present invention are defined within the body of the appended claims and the accompanying description and include at least the following: In a first aspect, according to claim 1, there is provided an audio personalization method for a first user. In another aspect, according to claim 2, there is provided a method for audio personalization for reference individuals. In another aspect, according to claim 15, there is provided an audio personalization system for a first user. - In another aspect, according to claim 16, there is provided an audio personalization system for a reference individual. [Brief description of the drawings]
[0005] A more complete understanding of the present disclosure and many of the attendant advantages will be readily obtained as the same becomes better understood by reference to the following detailed description, when considered in conjunction with the accompanying drawings, in which: [Figure 1] FIG. 1 is a schematic diagram of an entertainment device according to an embodiment of the present disclosure. [Figure 2A] 2A and 2B are schematic diagrams of audio characteristics relative to the head. [Figure 2B] 2A and 2B are schematic diagrams of audio characteristics relative to the head. [Figure 3A] 3A and 3B are schematic diagrams of the audio characteristics of the ear. [Figure 3B] 3A and 3B are schematic diagrams of the audio characteristics of the ear. [Figure 4A] 4A and 4B are schematic diagrams of an audio system used to generate data for calculation of head-related transfer functions according to embodiments herein. [Figure 4B] 4A and 4B are schematic diagrams of an audio system used to generate data for calculation of head-related transfer functions according to embodiments herein. [Diagram 5] FIG. 5 is a schematic diagram of the impulse responses of a user's left and right ear in the time and frequency domains. [Figure 6] FIG. 6 is a schematic diagram of head-related transfer function spectra for the left and right ears of a user. [Figure 7] FIG. 7 is a flow diagram of a method for audio personalization for a first user according to an embodiment of the present disclosure. [Figure 8] FIG. 8 is a flow diagram of a method for audio personalization for a reference individual according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0006] An audio personalization method and system is disclosed. In the following description, many specific details are presented to provide a thorough understanding of the embodiments of the present invention. However, it will be apparent to those skilled in the art that these specific details are not necessarily employed to implement the present invention. Conversely, specific details known to those skilled in the art are omitted where appropriate for clarity.
[0007] In an exemplary embodiment of the present invention, a system and / or platform suitable for implementing the methods and techniques herein may be an entertainment device such as a SONY PlayStation® 4 or 5 video game console.
[0008] For purposes of illustration, the following description is based on a PlayStation 4®, although it will be understood that this is a non-limiting example.
[0009] Referring now to the drawings, in which like reference numerals indicate identical or corresponding parts throughout the several views, Figure 1 illustrates generally the overall system architecture of a SONY® PlayStation4® entertainment device. A system unit 10 is provided along with a variety of peripherals that may be connected to the system unit.
[0010] The system unit 10 comprises an Accelerated Processing Unit (APU) 20, which is a single chip comprising a Central Processing Unit (CPU) 20A and a Graphics Processing Unit (GPU) 20B. The APU 20 has access to a Random Access Memory (RAM) unit 22.
[0011] APU 20 communicates with bus 40 optionally through I / O bridge 24 , which may be a separate component from APU 20 or may be part of APU 20 .
[0012] Also connected to the bus 40 are data storage components such as a hard disk drive 37 and a Blu-ray drive 36 operable to access data on a compatible optical disc 36A. Additionally, a RAM unit 22 may be in communication with the bus 40.
[0013] Optionally, an auxiliary processor 38 is also connected to bus 40. The auxiliary processor 38 may be provided to run or support an operating system.
[0014] The system unit 10 communicates with peripherals, as appropriate, via an audio / visual input port 31, an Ethernet port 32, a Bluetooth wireless link 33, a Wi-Fi wireless link 34, or one or more Universal Serial Bus (USB) ports 35. Audio and video may be output via an AV output 39, such as an HDMI port.
[0015] The peripherals may include monocular or stereoscopic video cameras 41, such as the PlayStation® Eye, stick-type video game controllers 42, such as the PlayStation® Move and conventional handheld video game controllers 43, such as the DualShock® 4, handheld entertainment devices 44, such as the PlayStation® Portable and PlayStation® Vita, keyboards 45 and / or mice 46, media controllers 47, e.g., in the form of a remote control, and headsets 48. Other peripherals, such as printers, or 3D printers (not shown), may be considered as well.
[0016] GPU 20B, optionally in cooperation with CPU 20A, generates video images and audio for output via AV output 39. Optionally, audio may be generated in cooperation with or instead by an audio processor (not shown).
[0017] The video and optionally the audio may be presented to a television 51. If supported by the television, the video may be stereoscopic. The audio may be presented to a home cinema system 52 in one of a number of formats, such as stereo, 5.1 surround sound, 7.1 surround sound, etc. The video and audio may similarly be presented to a head mounted display unit 53 worn by a user 60.
[0018] In operation, the entertainment device defaults to an operating system, such as one derived from FreeBSD® 9.0. The operating system may run on CPU 20A, auxiliary processor 38, or a mixture of the two. The operating system provides the user with a graphical user interface, such as the PlayStation® Dynamic Menu. The menus allow the user to access the operating system's functions and to select games and optionally other content.
[0019] When playing a game or any other content, a user typically receives audio from a stereo or surround sound system 52 or headphones when viewing the content on a stationary display 51, and from a stereo surround sound system 52 or headphones when viewing the content on a head mounted display ("HMD") 53.
[0020] In either case, the position of an in-game object relative to a static screen or the user's head position (or a combination of both) can be relatively easily displayed visually, but producing a corresponding audio effect is more difficult.
[0021] This is because an individual's perception of sound direction depends on the physical interaction with surrounding sounds caused by the physical properties of their head, but because every person's head is different, the physical interaction is unique.
[0022] Referring to FIG. 2A, one example of a physical interaction is interaural delay or time difference (ITD), which indicates the degree to which a sound is located to the left or right of the user (resulting in a relative change in arrival time at the left and right ears), which is a function of the listener's head size and face shape.
[0023] Similarly, referring to FIG. 2B, the interaural level difference (ILD) is associated with different loudness for the left and right ear and indicates the degree to which a sound is located to the left and right of the user (resulting in different degrees of attenuation due to the ears being relatively hidden from the sound source), again a function of head size and face shape.
[0024] In addition to this horizontal (left / right) discrimination, see FIG. 3A, the outer ear has asymmetric characteristics that vary from person to person and provide additional vertical discrimination to incident sounds; see FIG. 3B, small differences in path length between direct and reflected sounds due to these characteristics result in so-called spectral notches, which vary in frequency as a function of the height of the sound source.
[0025] Furthermore, these features are not independent: horizontal components such as ITD and ILD also vary as a function of source height because the shape of the face / head encountered by the sound waves on their way to the ears changes. Similarly, vertical components such as spectral notches also vary as a function of left / right position because the physical shape of the ear to the incoming sound and the resulting reflections also change with horizontal angle of incidence.
[0026] The result is a complex two-dimensional response in each ear that is a function of monaural cues, such as the spectral notch, and binaural or interaural cues, such as the ITD and ILD. The individual's brain learns to associate this response with the physical source of the object, and is able to distinguish between left / right, up / down, and indeed front / back to estimate the object's 3D location relative to the user's head.
[0027] It may be desirable to provide a user with sounds (e.g., via headphones) that replicate these characteristics, creating the illusion that objects in the game (or other sound sources in other consumable content) are at specific locations in space relative to the user, as in the real world. Such sounds are commonly known as binaural sound.
[0028] However, it will be appreciated that this is difficult to do without extensive testing, as each user is unique and requires a unique reproduction of characteristics.
[0029] In particular, it may be necessary to determine the user's in-ear response for multiple positions, for example in a surrounding sphere, and FIG. 4A shows a fixed speaker arrangement for this purpose, while FIG. 4B shows a simplified system in which, for example, the speaker device or the user can rotate in fixed increments, so that the speakers successively fill in the remaining sample points of the sphere.
[0030] 5, for a sound (e.g. an impulse such as a single delta or click) at each sample position, a recorded impulse response in the ear is obtained (e.g. using a microphone placed at the entrance to the ear canal) as shown in the graph above. The Fourier transform of these impulse responses yields the so-called head-related transfer function (HRTF) that describes for each ear the effect of the user's head on the received frequency spectrum for that point in space.
[0031] By measuring at many positions, the complete HRTF can be calculated, as shown partially for both the left and right ear in Figure 6 (frequency on the y-axis, azimuth on the x-axis). The brightness is a function of the Fourier transform value, and the dark areas correspond to the notches in the spectrum.
[0032] It will be appreciated that using a system such as that shown in FIGS. 4A and 4B to obtain HRTFs for each of the potential tens of millions of users of an entertainment device is impractical, as is supplying each individual user with some form of array system to perform self-testing.
[0033] Therefore, different techniques are disclosed in the embodiments herein.
[0034] In these embodiments, a system such as that shown in Figures 4A and 4B is used to obtain the complete HRTFs of a number of reference individuals to generate a library of HRTFs. This library may be small initially, for example, with individuals representative of several ages, ethnicities, and genders being tested, or simply a random selection of volunteers, beta testers, quality assurance testers, early adopters, etc. Over time, however, more individuals may be tested and the resulting HRTFs added to the library.
[0035] As with the HRTF test, each of these individuals performs a calibration test using, for example, an entertainment system and headphones described herein, or an HMD system (e.g., with headphones), or optionally a stereo or surround sound speaker system, and optionally two or more of these in series.
[0036] The calibration test asks the user to identify where in the space around them a sound is thought to be coming from. For a user wearing an HMD system, when a sound is played, the user can look in the direction the sound is thought to be coming from, and this direction can be measured (e.g., using head tracking and optionally eye tracking techniques known in the art). Alternatively or additionally, they can use one or more handheld controllers to move a reticle or other indicator to the expected location. In this latter case, they may move an indicator to a position on the screen that corresponds to where the sound is thought to be coming from, or, if the screen is displaying the user's notional position surrounded by a sphere or partial sphere, they can use the controller to move an indicator on the surface of the sphere to the notional location of the sound.
[0037] Alternatively or additionally, other means of input may be considered, such as camera-captured gesture input (e.g., pointing in the perceived direction from which a sound is coming).
[0038] Similarly, a location can be presented to the user graphically, and the user must control the location of the sound source to that location. In this case, pointing or other direct control would not be appropriate, as this does not require the user to estimate the location of the sound source. Rather, the sound source can be moved using, for example, joystick or joypad controls, or motion gestures (e.g., horizontal and / or vertical panning). However, this approach may be slower.
[0039] Thus, more generally, the user must try to match the presented sound to the presented location, either by controlling the location of the presented sound, or by controlling the location of the presented location.
[0040] Individuals whose complete HRTFs have been calculated and added to the library then perform this test (either localizing the sound or moving the sound to the localized location) using sounds transformed by a default HRTF (e.g. calculated using a dummy head) to generate a default binaural sound signal.
[0041] Depending on how an individual's morphology differs from the dummy's head morphology, the default HRTF used to drive binaural sound in headphones or speakers will differ in different ways from their natural HRTF, thus affecting their perception of where sound sources presented using the default HRTF actually are.
[0042] By testing multiple sound source locations in this way, the individual's location estimate (especially the degree of error in the location estimate) serves as a proxy description of how the individual's HRTFs differ from the default HRTFs. Such a proxy can also be thought of as a fingerprint of the complete HRTFs of a reference individual.
[0043] In embodiments herein, the home user may then perform the same calibration test. Optionally, if multiple types of audio distribution means are supported, rather than just headphones (and / or an HMD system where this is treated as equivalent to headphones), the user will indicate the type of audio system they are using (e.g., stereo or surround sound loudspeakers, or headphones, or an HMD system with headphones built in). This will affect the form of the default HRTF used (headphones, surround sound, etc.) and also the subset of proxy results of reference individuals in the library that are compared to the home user's results.
[0044] The home user can then perform the same calibration tests as the reference individual (either locating the sound for a set of locations or moving the sound to the identified locations) to estimate the location of the presented sound source using the default HRTF.
[0045] The closest pattern of location estimation errors in the set of proxy results is then taken to indicate the HRTF in the library that best matches the user's actual HRTF.
[0046] This indicated best matching HRTF can then be installed on the entertainment device as the HRTF for that user, thereby providing more realistic and accurate binaural sound to the user.
[0047] Additionally, the user's location estimates relative to the test sounds can be recorded, and if a new reference individual is added to the library, the user's location estimates can be tested against those of the new reference individual to see if they are a better match, for example, as a background service provided by a remote server. If a better match is found, the best-matching HRTF that is better represented can be installed as the HRTF for that user, thereby further improving the user's experience.
[0048] In this way, the HRTF of a user of an entertainment device can be estimated without, for example, placing a microphone in the user's ear canal or measuring an impulse response.
[0049] Advantageously, this allows potentially tens of millions of users to enjoy good quality binaural sound, and allows the sound quality to be improved as new reference individuals are added to the HRTF library.
[0050] The individuals selected to extend the library can also be carefully selected. For a representative set of reference individuals, a random distribution of users may be expected to map to each reference individual in approximately equal proportions. However, if a relatively large number of users map to a reference individual (e.g., exceeding a threshold variance in the number of users mapping to a reference individual), this indicates at least one of the following: i. The user population is not random (eg, due to demographics) so there are more people similar to this reference individual than is the norm. ii. Because the set of reference individuals is not sufficiently representative of the user, and there are gaps in the proxy result space surrounding this particular reference individual, people who are not actually very similar to this individual are mapped to them due to the lack of a better match.
[0051] In either case, it may be desirable to find other reference individuals morphologically similar to those currently in the library to provide more refined identification within this subgroup of the user population. Optionally, such individuals may be found by, for example, comparing face and profile (ears showing) photographs of the candidate individuals to help automatically assess head shape and ear shape. Such individuals may also be found using other methods, such as identifying individuals with similar demographics, or inviting close relatives or family members of existing individuals.
[0052] In this way, the HRTF library can optionally be augmented over time according to the characteristics of the user base.
[0053] If a suitable new reference individual cannot be found, or while a person is waiting to be added to the library, optionally, for a user who is close to two or more reference individuals but does not match any of them within a threshold, a blend of the HRTFs of the two or more reference individuals may be generated to provide a better estimate of the user's HRTF. This blend may be a weighted average or other combination according to the relative match (e.g., closeness in location error space to the vector of error values of the location estimation) of the HRTFs of the two or more reference individuals.
[0054] Optionally, as the library grows and as the user base grows, the library may be pre-filtered for a given user according to demographic criteria, for example according to one or more of age, sex, and ethnicity. This allows the set of reference individuals and therefore calibration test results for comparison to be reduced to a subset that matches these basic demographics. Then, only if the user's best match for location estimation still differs from those of the respective reference individual by a threshold, is that user compared to the full corpus of proxy results of the reference individuals. This may therefore reduce the computational overhead of the server making these comparisons, and also allow people who are not neatly located within their expected demographics (e.g., children with relatively large heads, or adults with relatively small heads) to find a good match within the broader library of reference individuals.
[0055] The above description assumes that a full calibration test is performed by a user at home, which may comprise localizing sounds at multiple positions, typically across a sphere or partial sphere surface, to capture the influence of the interconnected relationships between the horizontal and vertical audio features of the ITD, ILD, and spectral notches described above on the user's ability to estimate the location of an object whose sound has been processed using the default HRTF.
[0056] A full calibration test may be performed on a uniform grid of locations, or it may be performed with a non-linear distribution such that the density of test locations appears to spread out from the area in front of the user's resting line of sight to be sparsest behind it, for example prioritizing sounds within the user's normal field of view over sounds just outside it, then over sounds to the far left and right, then again over sounds behind the user.
[0057] A full calibration test may focus on areas where characteristics are known to vary, particularly if a large set of HRTFs of the type shown in Figure 6 were averaged (e.g., for reference individuals of a similar type, e.g., available based on age, sex, ethnicity, or other physiological measurements such as head size (or proxies such as hat size or perceived HMD wearing circumference)), it is conceivable that there would be a corresponding variance map that would indicate areas of the transfer functions of individuals that differ more than others, or in other words, where there is more room for discrimination in the calibration test.
[0058] As a result, there may be regions of space where reference individuals tend to exhibit larger estimation errors (e.g., variability above a threshold), and for these reference individuals, additional testing at nearby locations may provide useful additional differentiation between them.
[0059] Similarly, when a user is tested, if a large error exceeding such a threshold is identified, corresponding additional testing at nearby locations may be used to improve the results of the corresponding reference individual, and thus the selection of HRTFs. Furthermore, locations corresponding to large errors, or errors that appear to be outliers with respect to the candidate reference individual, may be reviewed to see if the errors are consistent and reproducible. If there is consistency, it may be retained and treated as important (e.g., to prompt the current user to add another reference individual, including the possibility of inviting the current user). If there is no consistency, the location may be fully or partially devalued when searching for the corresponding results of the reference individual.
[0060] In this way, the search space of the calibration test can be rapidly refined.
[0061] On the other hand, tests with a wide frequency range (e.g., bursts of white noise, or pops and bangs) can be useful for some characteristics (e.g., some notch measurements), while tests with a narrower frequency range can be useful for others. For example, pink noise below about 1.5 kHz can be useful for ITD-based estimation, while blue noise above 1.5 kHz can be more useful for ILD-based estimation. Other sounds such as chirps or pure tones may also be used as well, as well as natural sounds such as speech utterances, music, or environmental noise. Thus, a mix of wideband and narrowband sounds may be used for calibration to more clearly distinguish and characterize the impact of different aspects of a user's hearing on the location estimation.
[0062] Calibration tests typically randomize the selection of individual test locations within a predefined set of test locations so that neither the reference individual nor the home user learns audio position progression patterns.
[0063] However, it will be appreciated that a full calibration test may take a long time and may not be welcome or practical for a home user, but it will also be appreciated that the test can be performed incrementally, with additional test points adding to the user's proxy results and improving the potential accuracy of the match to the reference individual's proxy.
[0064] Thus, aspects of the tests can be prioritized or performed in priority order and refined with more data over any successive calibrations.
[0065] For example, measuring a height estimate of the centerline can provide an initial estimate of a height notch for the user's ear (or, more precisely, a pattern of location estimation errors characteristic of that notch). Similarly, measuring the horizontal position of the centerline can provide an initial estimate of the user's ITD and / or ILD (or, more precisely, a pattern of location estimation errors characteristic of these).
[0066] These test locations can again be randomized within just the vertical or horizontal ranges, or between both, or within a test set with a similar number of other predefined locations off these lines.
[0067] The user's results of this initial calibration test can be compared to corresponding initial results on a proxy for a reference individual to find the initial closest match, the corresponding HRTF that is again likely to provide a better experience for the user than the default.
[0068] The user can then enter a set of proxy results by revisiting the calibration test at a different time and continuing the test. The test locations can again prioritize specific locations that are more likely to provide specific discrimination for a given spectral notch, or provide ITD and / or ILD measurements over subsequent heights.
[0069] The user can run the calibration test again if desired - for example, a growing child may wish to do so annually as their head shape changes as they grow - similarly, an older individual may redo the calibration test if they suspect hearing loss in either ear.
[0070] Referring now also to Figures 7 and 8, in a summary embodiment herein, an audio personalization method for a reference individual therefore comprises the following steps:
[0071] In a first step s810, respective head-related transfer functions "HRTFs" are obtained for a corpus of reference individuals, as described elsewhere herein.
[0072] In a second step s820, each reference individual is tested with a calibration test. As described elsewhere herein, the calibration test typically comprises requesting each tested reference individual to match test sounds to test locations, either by controlling the location of the presented sounds or by controlling the location of the presented locations, for a sequence of test matches, as described elsewhere herein (e.g., by presenting a sequence of test sounds, which may be of the same type or may differ according to a predetermined scheme), where each test sound is presented at a location using a default head-related transfer function "HRTF"; receiving an estimate of each matching location from each tested reference individual, as described elsewhere herein (e.g., by receiving from the reference individual a respective location estimate for each test sound, or a final selected location of each sound estimated to match each test location); and calculating a respective location error for each estimate (e.g., the difference between the estimated location and the location of the sound, or the difference between the located sound source and the location), as described elsewhere herein, to generate a sequence of location estimation errors for each tested reference individual.
[0073] Then, in a third step s830, the sequence of location estimation errors for the reference individuals is associated with each obtained HRTF, as described elsewhere herein.
[0074] Meanwhile, in the summary embodiment of the present specification, an audio personalization method for a first user comprises the following steps.
[0075] A first step s710 comprises testing a first user with a calibration test, as described elsewhere herein.
[0076] The calibration test in turn comprises substep s712 of requesting the user to match the test sounds to test locations, either by controlling the location of the presented sounds or by controlling the location of the presented locations, for a sequence of test matches (e.g. by presenting a sequence of test sounds, which may again be of the same type or may differ according to a predetermined scheme), as described elsewhere herein, where each test sound is presented at a location using a default head-related transfer function "HRTF"; substep s714 of receiving an estimate of each matching location from the first user (e.g. by receiving from the first user a respective location estimate for each test sound, or a final selected location of each sound estimated to match each test location), as described elsewhere herein; and substep s716 of calculating a respective error for each estimate (e.g. the difference between the user-estimated location and the location of the sound, or the difference between the sound source and the location positioned by the user), as described elsewhere herein, to generate a sequence of location estimation errors for the first user.
[0077] Next, a second step s720 comprises comparing at least some of the first user's location estimation errors with estimation errors of the same locations previously generated for at least a subset of the corpus of reference individuals, as previously described in this specification.
[0078] Then, a third step s730 comprises identifying the reference individual whose compared location estimation error most closely matches that of the first user, as previously described herein.
[0079] Then, a fourth step s740 comprises applying the HRTFs previously obtained for the identified reference individual to the first user, as previously described herein.
[0080] It will be appreciated that typically the method relating to the reference individual will be performed by a provider of a video game console or other content playback device, or a provider of system software for such a console or device, or a provider of audio toolkits for software developers of such a console or device, while the method relating to the first user will be performed on behalf of the first user using his or her own console or other content playback device.
[0081] As a result, although the methods can be employed independently, the method with respect to the first user presupposes that the method with respect to the reference individuals has been implemented to the extent that at least some set of HRTFs and location estimation errors for some reference individuals exists.
[0082] However, it will be appreciated that the two methods can also be considered as part of a single broader method, such as mass user audio composition.
[0083] As will be apparent to one of ordinary skill in the art, variations of the above methods that correspond to the operation of various embodiments of the methods and / or apparatus described herein and in the claims are considered to be within the scope of this disclosure, including, but not limited to, the following: - Re-compare users to the corpus from time to time as the corpus grows, as described elsewhere herein. Thus, when a predetermined number of reference individuals with available sequences of HRTFs and associated location estimation errors have been added to the corpus, compare at least some of the location estimation errors of the first user to the estimation errors of the same location for at least a subset of the corpus of additional reference individuals. Then, as described elsewhere herein, use the HRTFs obtained for the additional reference individuals for the first user if the additional reference individuals have a closer match in the compared location estimation errors to those of the first user than the currently identified reference individuals. As described elsewhere herein, the subset of the corpus is selected as a function of demographic details of the first user and a reference individual. - As described elsewhere herein, each location comprises at least a subset of locations selected due to at least a threshold variance in location estimation errors for a subset of reference individuals. As described elsewhere herein, each sound used in the calibration test comprises one or more sounds selected from the list consisting of narrowband sounds, broadband sounds, impulse sounds, tones, chirps, and voices. As described elsewhere herein, for calibration testing, the locations are selected in a predetermined series of subsets from a set of predetermined locations. As described elsewhere in this specification, in this case, optionally, the subset comprising locations on the horizontal centerline and the subset comprising locations on the vertical centerline are included within the first N subsets of a predetermined set of subsets, where N is between 2 and 5. o As described elsewhere in this specification, again optionally, the comparing step s720, determining step s730 and using step s740 are performed after a predetermined number of subsets have been completed within a predetermined series of subsets. As described elsewhere herein, in this case, optionally, the comparing, characterizing and using steps are performed again when the first user subsequently undergoes a calibration test with a predetermined number of subsequent subsets of the predetermined series of subsets. As described elsewhere herein, for calibration testing, each location is randomly selected from at least a subset of the predetermined locations (which may comprise one or more subsets from a predetermined set of subsets). - As described elsewhere herein, if a first reference individual is identified as the best match for a user by a threshold number of more than other reference individuals, additional reference individuals are selected that have morphological similarity to the first reference individual within a predetermined tolerance. As described elsewhere herein, if there is no single reference individual whose compared location estimation error matches that of the first user within a predetermined threshold level of match, the method comprises blending the HRTFs of the closest M matching reference individuals, where M is a value greater than or equal to 2, and using the blended HRTF for the first user.
[0084] It will be appreciated that the above methods may be carried out on conventional hardware suitably adapted to be applicable by software instructions or by the inclusion or substitution of dedicated hardware.
[0085] Thus, the necessary adaptations to existing parts of the conventional equivalent devices may be implemented in the form of a computer program product including processor-implementable instructions stored on a non-transitory machine-readable medium such as a floppy disk, optical disk, hard disk, solid state disk, PROM, RAM, flash memory, or any combination of these or other storage media, or may be realized in hardware as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array) or other configurable circuitry suitable for use in adapting the conventional equivalent devices. Alternatively, such computer programs may be transmitted via data signals over a network such as an Ethernet, a wireless network, the Internet, or any combination of these or other networks.
[0086] While the data required to calculate the HRTFs may be specialist equipment as shown in Figures 4A and 4B, the device used to perform the calibration tests and carry out steps such as relating location estimation errors to individuals and / or HRTFs, comparing the results, identifying the best match and using the corresponding HRTFs may be a video game console such as a PS4 (registered trademark) or PS5 (registered trademark), or equivalent development kit, a PC, etc.
[0087] Thus, in a summarized embodiment, an audio personalization system for a first user includes: a test processor (e.g., CPU 20A) configured (e.g., by suitable software instructions) to test a first user with a calibration test, the calibration test comprising: requesting the user to match test sounds to test locations, either by controlling the location of the presented sounds or by controlling the location of the presented locations, for a sequence of test matches (e.g., by presenting a sequence of test sounds, which again may be of the same type or may differ according to a predetermined scheme), as described elsewhere herein, where each test sound is presented at a location using a default head-related transfer function "HRTF"; receiving estimates of each matching location from the first user (e.g., by receiving from the first user a respective location estimate for each test sound), as described elsewhere herein; and calculating a respective error for each estimate, generating a sequence of location estimation errors for the first user, as described elsewhere herein; a comparison processor (e.g., CPU 20A) configured (e.g., by suitable software instructions) to compare at least some of the first user's location estimation errors with previously generated estimation errors of the same location for at least a subset of the corpus of reference individuals, as described elsewhere herein, the comparison processor also configured (e.g., by suitable software instructions) to identify the reference individual whose compared location estimation errors most closely match those of the first user; an HRTF processor (e.g., CPU 20A) configured (e.g., by suitable software instructions) to use previously obtained HRTFs for the identified reference individual for the first user, as described elsewhere herein; The entertainment device 10 may include:
[0088] For example, it will be appreciated that the responsibilities of the comparison processor may be divided between the entertainment device and a remote server that also maintains the location estimation errors for the corpus of reference individuals. Thus, within the entertainment device, the comparison processor is configured to perform the comparison, which may be performed either locally (e.g., by performing the comparison) or remotely (e.g., by sending the first user's location estimation error to the server and requesting a comparison).
[0089] Likewise, it will be appreciated that the HRTF processor may receive appropriate HRTF data from such a remote server.
[0090] Similarly, in a summary embodiment, an audio personalization system for a reference individual includes: a storage (such as a HDD 37 in conjunction with the CPU 20A) configured (e.g., by suitable software instructions) to store respective head-related transfer functions "HRTFs" for the corpus of reference individuals; a test processor (e.g., CPU 20A) configured (e.g., by suitable software instructions) to test each reference individual with a calibration test, the calibration test comprising: requesting the reference individual to match the test sounds to test locations, either by controlling the location of the presented sounds or by controlling the location of the presented locations, for a sequence of test matches (e.g., by presenting a sequence of test sounds, which may be of the same type or which may differ according to a predetermined scheme), as described elsewhere herein, where each test sound is presented at a location using a default head-related transfer function "HRTF"; receiving an estimate of each matching location from each tested reference individual, as described elsewhere herein (e.g., by receiving from the reference individual a respective location estimate for each test sound, or a final selected location of each sound estimated to match each test location); and calculating a respective location error for each estimate, as described elsewhere herein, to generate a sequence of location estimation errors for each tested reference individual; an association processor (e.g., CPU 20A) configured (e.g., by suitable software instructions) to associate the sequence of location estimation errors for the reference individuals with each acquired HRTF; The entertainment device 10 may also be a development kit or a server.
[0091] It will be appreciated that the calibration tests for the first user and the reference individual are typically the same.
[0092] The foregoing discussion discloses and describes merely exemplary embodiments of the present invention. As will be understood by those skilled in the art, the present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. Accordingly, the disclosure of the present invention, as well as the other claims, are intended to be illustrative but not limiting of the scope of the present invention. The present disclosure, including any readily identifiable variations of the teachings herein, defines in part the scope of the preceding claim terms, so that the inventive subject matter is not dedicated to the public.
Claims
1. 1. A method of audio personalization for a first user, comprising: testing the first user with a calibration test, said calibration test comprising: requiring the user to match test sounds to test locations for a sequence of test matches, either by controlling the location of the presented sound or by controlling the location of the presented location, where each test sound is presented at a location using a default head-related transfer function "HRTF"; receiving an estimate of each matching location from the first user; and calculating a respective error for each estimate to generate a sequence of location estimation errors for the first user; comparing at least some of the location estimation errors of the first user with previously generated estimation errors of the same locations for at least a subset of a corpus of reference individuals; identifying a reference individual whose compared location estimation error most closely matches that of the first user; using previously obtained HRTFs for the identified reference individual for the first user; An audio personalization method comprising:
2. 1. A method for audio personalization for a reference individual, comprising: obtaining respective head-related transfer functions "HRTFs" for a corpus of reference individuals; testing each reference individual with a calibration test, said calibration test comprising: requiring the reference individual to match test sounds to test locations for a sequence of test matches, either by controlling the location of the presented sound or by controlling the location of the presented location, where each test sound is presented at a location using a default head-related transfer function "HRTF"; receiving an estimate of each matching location from said reference individual; and calculating a respective error for each estimation to generate a sequence of location estimation errors for each tested reference individual; associating the sequence of location estimation errors of the reference individuals with each acquired HRTF; An audio personalization method comprising:
3. When a predetermined number of reference individuals for whom sequences of HRTFs and associated location estimation errors are available are added to the corpus, comparing at least some of the location estimation errors of the first user to the estimation errors of the same locations for at least a subset of the corpus of additional reference individuals; if the additional reference individual has a compared location estimation error that is a closer match to that of the first user than the currently identified reference individual; using the HRTFs obtained for the additional reference users for the first user; The audio personalization method of claim 1.
4. the subset of the corpus is selected as a function of demographic details of the first user and the reference individual.
4. The audio personalization method according to claim 1 or 3.
5. For the calibration test, each location is selected in a predetermined series of subsets from a predetermined set of locations; comparing at least some of the first user's location estimation errors to corresponding estimation errors for at least a subset of the corpus of reference individuals; identifying a reference individual whose compared location estimation error most closely matches that of the first user; using the HRTFs obtained for the identified reference user for the first user, the step being performed after a predetermined number of subsets have been completed within the predetermined series of subsets.
5. A method for audio personalization according to claim 1, 3 or 4.
6. when the first user subsequently undergoes the calibration test with a predetermined number of subsequent subsets of the predetermined series of subsets, the comparing, identifying and using steps are performed again.
6. The audio personalization method of claim 5.
7. If there is no single reference individual whose compared location estimation error matches that of the first user within a predetermined threshold level of agreement, the method further comprises: blending the HRTFs of the M closest matching reference individuals, where M is a value equal to or greater than 1; using the blended HRTFs for the first user; 7. The method of claim 1, 3 or 6, comprising:
8. For the calibration test, Each location is selected from a predetermined set of locations in a predetermined set of subsets.
5. An audio personalization method according to any one of claims 1 to 4.
9. The subsets having locations on the horizontal centerline and the subsets having locations on the vertical centerline are included within a first N subsets of the predetermined set of subsets, where N is between 2 and 5.
9. The audio personalization method of claim 8.
10. each location comprises at least a subset of locations selected by having at least a threshold variance in location estimation errors for a subset of reference individuals; 10. An audio personalization method according to any one of claims 1 to 9.
11. Each sound used in the calibration test is i. Narrowband sounds; ii. Broadband sound; iii. Impulse sound; iv. Tone, v. chirps, and vi. Voice; and one or more selected from the list consisting of:
11. An audio personalization method according to any one of claims 1 to 10.
12. For calibration tests, Each location is randomly selected from at least a subset of the predetermined locations.
12. An audio personalization method according to any one of claims 1 to 11.
13. If a first reference individual is identified as the best match for the user by a threshold number of times more than other reference individuals, an additional reference individual is selected that has a morphological similarity to the first reference individual within a predetermined tolerance; 13. An audio personalization method according to any one of claims 1 to 12.
14. A computer program comprising computer executable instructions adapted to cause a computer system to carry out the method according to any of claims 1 to 13.
15. 1. An audio personalization system for a first user, comprising: a test processor configured to test a first user with a calibration test, the calibration test comprising: requiring the user to match test sounds to test locations for a sequence of test matches, either by controlling the location of the presented sound or by controlling the location of the presented location, where each test sound is presented at a location using a default head-related transfer function "HRTF"; receiving an estimate of each matching location from the first user; and a test processor for calculating a respective error for each estimate to generate a sequence of location estimation errors for the first user; a comparison processor configured to compare at least some of the location estimation errors of the first user with estimation errors of the same locations previously generated for at least a subset of a corpus of reference individuals, the comparison processor configured to identify a reference individual whose compared location estimation errors most closely match those of the first user; a HRTF processor configured to use previously obtained HRTFs for the identified reference individual for the first user; An audio personalization system comprising:
16. 1. An audio personalization system for a reference individual, comprising: a storage configured to store respective head-related transfer functions "HRTFs" for the corpus of reference individuals; a test processor configured to test each reference individual with a calibration test, said calibration test comprising: requiring the reference individual to match test sounds to test locations in a sequence of test matches, either by controlling the location of the presented sound or by controlling the location of the presented location, where each test sound is presented at a location using a default head-related transfer function "HRTF"; receiving an estimate of each matching location from said reference individual; and a test processor for calculating a respective error for each estimate to generate a sequence of location estimation errors for each tested reference individual; an association processor configured to associate the sequence of location estimation errors of the reference individuals with each obtained HRTF; An audio personalization system comprising:
Citation Information
Patent Citations
Three-dimensional sound field reproduction system
JP2011182135A
HRTF measurement method, HRTF measurement device, and program
WO2018110269A1
Acoustic processing device, acoustic processing method, and acoustic processing program
WO2020189263A1