Spatial Repositioning of Multiple Acoustic Streams
Personalized BRIRs in headphones allow seamless playback of music and clear differentiation of incoming calls by positioning sound sources at distinct spatial acoustic positions, enhancing audio immersion.
Patent Information
- Application Number
- JP2019221087
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-12-07
- Filing Date
- 2019-12-06
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2039-12-06
AI Technical Summary
Existing systems fail to seamlessly play music and distinguish incoming calls without interruption, and conventional methods for positioning acoustic streams in headphones are ineffective due to difficulty in localizing sound sources accurately.
Utilizing personalized binaural room impulse responses (BRIRs) to position music and voice calls at distinct spatial acoustic positions, enabling immersive audio experiences by directing sound sources to foreground and background positions based on individual listener characteristics.
Enables continuous playback of music and clear differentiation between incoming calls by positioning sound sources effectively, providing an immersive audio experience through personalized spatial acoustic rendering.
Smart Images

Figure 0007705647000001 
Figure 0007705647000002 
Figure 0007705647000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the benefit of priority of U.S. Patent Application No. 62 / 614,482, filed on January 7, 2018, "METHOD FOR GENERATING CUSTOMIZED SPATIAL AUDIO WITH HEAD TRACKING", Singapore Patent Application No. 10201510822Y, filed on December 31, 2015, "A METHOD FOR GENERATING A CUSTOMIZED / PERSONALIZED HEAD RELATED TRANSFER FUNCTION", and International Patent Application No. PCT / SG2016 / 050621, filed on December 28, 2016, "A METHOD FOR GENERATING A CUSTOMIZED / PERSONALIZED HEAD RELATED TRANSFER FUNCTION", and incorporates by reference the entire disclosure content of all of them, and incorporates all of its content herein. Further, this application incorporates by reference the entire disclosure content of U.S. Patent Application No. 15 / 969,767, filed on May 2, 2018, "SYSTEM AND A PROCESSING METHOD FOR CUSTOMIZING AUDIO EXPERIENCE" and U.S. Patent Application No. 16 / 136,211, filed on September 19, 2018, "METHOD FOR GENERATING CUSTOMIZED SPATIAL AUDIO WITH HEAD TRACKING".
[0002] The present invention relates to a method and system for generating sound for rendering via headphones. More particularly, the present invention relates to using a database of personalized spatial acoustic transfer functions having room impulse response information associated with spatial acoustic positions in conjunction with an acoustic stream, and generating a more realistic acoustic rendering via headphones by generating spatial acoustic positions using the personalized spatial acoustic transfer functions.
Background Art
[0003] Users often listen to music when the phone rings and may want to continue listening without interrupting the music. Unfortunately, most phones are configured to mute the music when receiving an incoming call. Therefore, there is a need for an improved system that can continue to play sounds such as music without interruption when receiving an incoming call and that allows the user to distinguish between two different sound sources. SUMMARY OF THE INVENTION
[0004] To achieve the above, in various embodiments, the present invention provides a processor system configured to provide a binaural signal to headphones, the system comprising means for arranging sound in a first input sound channel at a first position such as a foreground position, and means for arranging sound in a second input sound channel at a second position such as a background position.
[0005] In some of the embodiments of the present invention, the system includes a database of personalized spatial acoustic transfer functions having room impulse response information (such as HRTF or BRIR, etc.) associated with spatial acoustic positions in conjunction with at least two acoustic streams. In combination with this, by using personalized BRIRs for at least two locations in conjunction with two input acoustic streams, foreground and background spatial acoustic sources are established so that the listener can obtain an immersive experience through the headphones. BRIEF DESCRIPTION OF THE DRAWINGS
[0006]
Figure 1
Figure 2
Figure 3
Best Mode for Carrying Out the Invention
[0007] Hereinafter, preferred embodiments of the present invention will be referred to in detail. Examples of preferred embodiments are shown in the accompanying drawings. The present invention will be described in relation to these preferred embodiments, but it should be understood that the present invention is not intended to be limited to such preferred embodiments. Rather, it is intended to cover alternatives, improvements, and equivalents that can be included within the spirit and scope of the present invention as defined by the appended claims. In the following description, many specific details are shown to enable a thorough understanding of the present invention. The present invention can be practiced without some or all of these specific details. In other instances, well-known mechanisms have not been described in detail so as not to unnecessarily obscure the present invention.
[0008] It should be noted that throughout this specification, the same numbers represent the same parts throughout the various drawings. The various drawings illustrated and described in this specification are used to show the various features of the present invention. Unless a particular feature is shown in one drawing and not in another, and there is no specific designation or essential structural incorporation prohibition of the feature, it is understood that these features can be adapted to be included in the embodiments shown in other drawings as if they were fully illustrated. Unless otherwise specified, the drawings are not necessarily drawn to scale. Any dimensions on the drawings are not intended to limit the scope of the present invention and are merely examples.
[0009] Binaural technology generally refers to technology related to or used in both ears and enables a user to recognize sound in a three-dimensional space. This is achieved, in some embodiments, by the determination and use of a binaural room impulse response (BRIR) and its related binaural room transfer function (BRTF). The BRIR simulates the interaction of sound waves from a speaker with the listener's ears, head, and torso, as well as the room's walls and other objects. Alternatively, in some embodiments, a head-related transfer function (HRTF) is used. The HRTF is the transfer function in the frequency domain corresponding to the impulse response representing the interaction in a free-field environment. That is, the impulse response here represents the sound interaction with the listener's ears, head, and torso.
[0010] According to known methods for determining the HRTF or BRTF, a real or dummy head microphone and a binaural microphone are used to record the stereo impulse response (IR) for each of several speaker positions in an actual room. That is, for each position, a pair of impulse responses is generated, one for each ear. This pair is referred to as the BRIR. Then, these BRIRs can be used to convolve (filter) a music track or other audio stream and mix the results for playback via headphones. When the correct equalization is applied, the music channels will be heard as if they were being played back at the speaker position in the room where the BRIR was recorded.
[0011] Users often listen to music when the phone rings and may want to continue listening without interrupting the music when receiving an incoming call. Instead of invoking a mute function, two separate acoustic signals, namely the phone and the music, can be supplied to the same channel. However, for humans, it is generally difficult to distinguish sound sources coming from the same direction. To solve this problem, according to one embodiment, when the phone rings, the music is directed from a first position to a speaker or channel at a second position such as a background position. That is, the music and the voice communication are arranged at different positions. Unfortunately, the method of positioning these rendering acoustic streams, while enabling separation of sound sources when used in combination with a multi-speaker setup, most of today's voice communications are via mobile phones, and these are typically not connected to a multi-channel speaker setup. Furthermore, even when such a method is used in combination with a multi-channel setup, when the acoustic source is specified by panning for a position that does not exactly match the physical position of the speaker, optimal results cannot be obtained. This is partly due to the fact that it is difficult for the listener to precisely localize the spatial acoustic position when approximating the position by the conventional panning method that moves the perceived acoustic position to a location between the multi-channel speaker positions.
[0012] The present invention solves the problem of voice communication via headphones by automatically positioning voice calls and music at different spatial acoustic positions by using virtualized positions with transfer functions that simulate at least the effects of an individual's head, torso, and ears on sound, such as by using HRTF. More preferably, by processing the acoustic stream with BRIR, the indoor effects on sound are taken into account. However, off-the-shelf, non-personalized BRIR datasets give most users a poor sense of directionality and a poor sense of distance to the perceived sound source. This may cause difficulties in distinguishing sound sources.
[0013] To solve these further problems, in some embodiments of the present invention, personalized BRIRs are used. In one embodiment, a personalized HRTF or BRIR dataset is generated by inserting a microphone into the listener's ear and recording the impulse response in a recording session. This is a time-consuming process and may be inconvenient to include in the sale of a mobile phone or other acoustic unit. In another embodiment, the sound sources of voice and music are localized to two separate locations, a first (e.g., foreground) and a second (e.g., background), for an individual listener by using personalized BRIRs (or related BRTFs) derived from the extraction of image-based characteristics. The characteristics are used to determine an appropriate personalized BRIR from a database having a pool of candidate personalized spatial acoustic transfer functions for a plurality of individuals being measured. The personalized BRIRs corresponding to at least two separate spatial acoustic positions are preferably used to direct the first and second acoustic streams to two different spatial acoustic positions.
[0014] Furthermore, since it is known that humans are better able to distinguish between two sound sources when one of the two sound sources is determined by the listener to be closer and the other of the two sound sources is determined to be farther away, in some embodiments, the personalized BRIRs derived using the extracted image-based characteristics are used to automatically place the music at a distance in the background spatial position and the voice at a closer distance.
[0015] In another embodiment, the extracted image-based characteristics are generated by a mobile phone. In another embodiment, when it is determined that the priority of the voice call is low and a control signal from the listener is received, for example, by activating a switch, the voice call is directed to the foreground / background and the music is directed to the foreground. In yet another embodiment, when it is determined that the priority of the voice call is low and a control signal from the listener is received, the apparent distance of the voice call is increased and the apparent distance of the music is decreased using a personalized BRIR corresponding to different distances in the same direction.
[0016] It should be understood that most of the embodiments described herein describe a personalized BRIR for use with headphones, but the techniques for positioning media streams in conjunction with the described voice communications are also extensible to any suitable transfer function customized for the user according to the steps described with respect to FIG. 3.
[0017] It is understood that the scope of the present invention is intended to cover placing each of the first acoustic source and the voice communication at any position around the user. Further, the foreground and background used herein are not intended to be limited to the areas in front of or behind the listener. Rather, the foreground should be interpreted in its most general sense as representing the more prominent or important of two distinct positions, while the background represents the less prominent of the distinct positions. Further, it should be noted that the scope of the present invention, in its most general sense, is to direct a first acoustic stream to a first location and a second acoustic stream to a second spatial acoustic location using HRTF or BRIR in accordance with the techniques described herein. Further, some embodiments of the present invention can be extended to the selection of any directional position around the user for either the foreground position or the background position by the simultaneous application of signal attenuation, instead of assigning a near distance to the foreground position and a far distance to the background position. First, a filtering circuit representing the foreground position and the background position by the application of a pair of BRIRs according to an embodiment of the present invention is shown in its simplest form below.
[0018] FIG. 1 is a diagram showing the spatial acoustic position of processed sound according to some embodiments of the present invention. First, listener 105 can listen to a first acoustic signal such as music through headphones 103. Using the BRIR applied to the first acoustic stream, the listener perceives that the first acoustic stream is arriving from the first acoustic position 102. In some embodiments, this is the foreground position. In one embodiment, a technique places this foreground position at the 0° position relative to listener 105. When a triggering event such as an incoming phone call occurs in one embodiment, the first acoustic signal is guided to the second position 104, while the second stream (e.g., voice communication or phone call) is guided to the first position (102). In the illustrated exemplary embodiment, this second position is arranged at the 200° position and in some embodiments is described as an unobtrusive or background position. The 200° position is merely selected as a non-limiting example. The arrangement of the acoustic stream at this second position is preferably realized using a BRIR (or, BRTF) corresponding to the azimuth, elevation, and distance of the second position of the target listener.
[0019] In one embodiment, the transition of the first acoustic stream to the second position (e.g., background) occurs abruptly without giving any sense that the first acoustic stream is moving through an intermediate spatial position. This is illustrated by path 110 which does not show an intermediate spatial position. In another embodiment, the sound is positioned at intermediate points 112 and 114 for a short transition period, giving a sense of moving directly or alternatively in an arc from the foreground position 102 to the background position 104. In a preferred embodiment, the BRIRs for the intermediate points 112 and 114 are used to spatially position the acoustic stream. In an alternative embodiment, the sense of movement is achieved by using the BRIRs for the foreground and background positions and panning between the virtual speakers corresponding to these foreground and background positions. In some embodiments, the user can recognize that a voice communication (e.g., a phone call) is not worthy of the priority status and choose to demote the phone to a second position (e.g., the background position) or a third position selected by the user and return the music to the first (e.g., foreground) position. In one embodiment, this is performed by sending back the acoustic stream corresponding to the music to the foreground (first) position 102 and sending the voice communication to the background position 104. In another embodiment, this re-prioritization is performed by moving the voice call away from the listener's head 105 and bringing the music closer. This is preferably done by assigning a new HRTF or BRTF for the listener, which is captured at different distances and represents a new distance by calculation or interpolation from the captured measurements. For example, to increase the priority of the music from the background position 104, the apparent distance can be shortened to the spatial acoustic position 118 or 116. Such a shortening of the distance is preferably achieved by processing the acoustic stream of the music with a new HRTF or BRTF, which results in an increase in the volume of the music relative to the voice communication signal.In some embodiments, again, the distance of the audio signal from the listener's head 105 can be simultaneously increased by the selection or interpolation of the captured HRTF / BRTF values. This interpolation / calculation can be performed using three or more points. For example, to obtain a point that is the intersection of two lines (AB and CD), the interpolation / calculation may require points A, B, C, and D.
[0020] Alternatively, the spatial acoustic position for generating the audio communication can be maintained at a fixed position or increased in the re-grading step. In some embodiments, two separate acoustic streams enjoy equal prominence.
[0021] In yet other embodiments, the user can select a spatial acoustic position for at least one of the above streams from the user interface, and more preferably, can select a single or multiple locations for all of the above streams.
[0022] FIG. 2 is a diagram showing a system for simulating acoustic sources and voice communications at different spatial acoustic positions according to some embodiments of the present invention. FIG. 2 roughly shows two different streams (202 and 204) entering the spatial acoustic positioning system by using separate filter pairs (i.e., filters 207, 208) for the first spatial acoustic position and filters 209, 210 for the second spatial acoustic position. For all filtered streams, gains 222-225 can be applied before the signal for the left cup of the headphones is added by adder 214 and the filtered result for the right cup of headphones 216 is similarly added by adder 215. This group of hardware modules shows the underlying principles involved, but other embodiments use BRRIs or HRTFs stored in a memory such as memory 732 of an acoustic rendering module 730 (such as a mobile phone), as shown in FIG. 3. In some embodiments, the listener is assisted in identifying these spatial acoustic positions by the fact that the first and second spatial acoustic positions are generated by selecting a transfer function having an indoor response in addition to the individual's HRTF. In a preferred embodiment, the first and second positions are determined using BRIRs customized for the listener.
[0023] Systems and methods for rendering via headphones work best when the HRTF or BRTF is individualized for the listener by direct in-ear microphone measurements or a personalized BRIR / HRIR dataset when in-ear microphone measurements are not used. According to a preferred embodiment of the present invention, a custom method for generating the BRIR is used, which, as generally shown in FIG. 3, includes the extraction of image-based characteristics from the user and the determination of the appropriate BRIR from a BRIR candidate pool. More specifically, FIG. 3 shows a system according to an embodiment of the present invention for generating a customized HRTF for customization, obtaining customized listener characteristics, selecting the customized HRTF of the listener, providing a rotation filter adapted to function correctly with relative movement of the user's head, and rendering acoustics modified by the BRIR. The extraction device 702 is a device configured to identify and extract the acoustic-related physical characteristics of the listener. In a preferred embodiment, block 702 can be configured to directly measure these characteristics (e.g., ear height), but the relevant measurement results are extracted from an image of the user obtained so as to include at least one or both ears of the user. Although the processing required for the extraction of these characteristics is preferably performed in the extraction device 702, it may be performed elsewhere. As a non-limiting example, these characteristics can also be extracted by a processor of the remote server 710 after receiving the image from the image sensor 704.
[0024] In a preferred embodiment, the image sensor 704 acquires an image of the user's ear, and the processor 706 is configured to extract relevant characteristics of the user and transmit them to the remote server 710. For example, in one embodiment, the use of a dynamic shape model is used to identify landmarks in the auricle image, and using these landmarks, their respective geometric relationships, and the straight-line distances, the characteristics of the user related to the generation of a customized BRIR from a set of stored BRIR data sets, i.e., the candidate pool of BRIR data sets, can be identified. In other embodiments, the RGT model (regression tree model) is used to extract the characteristics. In still other embodiments, the characteristics are extracted by using machine learning such as neural networks and other forms of artificial intelligence (AI). An example of a neural network is a convolutional neural network. Details of a plurality of methods for identifying the unique physical characteristics of a new listener are described in International Patent Application No. PCT / SG2016 / 050621, "A Method for Generating a customized Personalized Head Related Transfer Function", filed on December 28, 2016, the entire disclosure of which is incorporated herein by reference.
[0025] The remote server 710 is preferably accessible via a network such as the Internet. The remote server preferably comprises a selection processor 710 that accesses the memory 714 and determines the most matching BRIR data set using the physical characteristics or other image-related characteristics extracted by the extraction device 702. The selection processor 712 preferably accesses a memory 714 having a plurality of BRIR data sets. That is, for the azimuth and elevation angles, and perhaps also for the head tilt, preferably for each point at an appropriate angle, each data set in the candidate pool will have a BRIR pair. For example, by obtaining measurement results every 3° for the azimuth and elevation angles, a BRIR data set of a sampling individual that constitutes a BRIR candidate pool can be generated.
[0026] As described above, although these are preferably derived from measurements using an in-ear microphone for a medium-sized (i.e., over 100 people) group, they can function correctly even for smaller groups of individuals and are stored with similar image-related characteristics associated with each BRIR set. Some of these are generated directly by measurement and some are generated by interpolation, and a spherical grid of BRIR pairs can be constructed. Even for a partially measured / partially interpolated grid, if the appropriate BRIR pair for a point from the BRIR data set is identified by the appropriate azimuth and elevation values, interpolation is also possible for another point not located on the grid line. For example, any suitable interpolation method can be used, preferably in the frequency domain, including but not limited to adjacent linear interpolation, bilinear interpolation, and spherical triangular interpolation.
[0027] In one embodiment, each BRIR data set stored in the memory 714 includes at least a global grid of the listener. In such a case, any angle of azimuth or elevation (on the horizontal plane around the listener, i.e., at ear height) can be selected with respect to the source placement. In other embodiments, the BRIR data set is more limited, and in one example, it is limited to the BRIR pairs necessary for generating a speaker placement in a room that matches a conventional stereo placement (i.e., +30° and -30° with respect to the zero position straight ahead, or a speaker placement for a multi-channel placement not limited to a 5.1 system or a 7.1 system, etc., in another subset of the global grid).
[0028] HRIR is the head impulse response. This completely describes the propagation of sound from the sound source to the listener in the time domain under anechoic conditions. Most of the information contained in this relates to the physiological functions and anthropometry of the person being measured. HRTF is the head-related transfer function. This is the same as HRIR except that it is a description in the frequency domain. BRIR is the binaural room impulse response. This is the same as HRIR except that since it is measured in a room, it additionally includes the captured room response for a specific configuration. BRTF is the frequency domain version of BRIR. In this specification, since BRIR can be easily replaced with BRTF and similarly HRIR can be easily replaced with HRTF, it is understood that the embodiments of the present invention are intended to cover these easily replaceable steps even if they are not specifically described. For this reason, for example, if the description represents access to another BRIR dataset, it is understood that access to another BRTF is covered.
[0029] FIG. 3 further shows the logical relationship of samples for data stored in the memory. The memory is shown as including a plurality of individual BRIR data sets (e.g., HRTF DS1A, HRTF DS2A, etc.) in column 716. These are indexed and accessed by characteristics associated with each BRIR data set, preferably image-related characteristics. The related characteristics shown in column 715 can match the identification of a new listener with the characteristics associated with the BRIRs measured and stored in columns 716, 717, and 718. That is, they act as indexes to a candidate pool of BRIR data sets shown in these columns. Column 717 represents the BRIR stored at the reference position zero, which is associated with the rest of the BRIR data set, enabling efficient storage and processing by combining with a rotation filter during monitoring of the listener's head rotation and its corresponding response. Details of this option are described in detail in co-pending application Ser. No. 16 / 136,211, "METHOD FOR GENERATING CUSTOMIZED SPATIAL AUDIO WITH HEAD TRACKING," filed Sep. 19, 2018, the entire contents of which are incorporated herein by reference.
[0030] Generally, one purpose of accessing a candidate pool of BRIR (or HRTF) datasets is to generate acoustical response characteristics (such as a BRIR dataset) customized for a person. In some embodiments, as described above, these are used to process and position input acoustic signals such as voice communications and media streams to accurately recognize the spatial acoustics associated with a first position and a second position. In some embodiments, generating customized acoustical response characteristics such as personalized BRIR includes extracting image-related characteristics such as an individual's biometric data. For example, this biometric data can include data related to the auricle, generally the person's ear, head, and / or shoulders. In another embodiment, intermediate datasets are generated by using processing methods such as (1) multiple match, (2) multiple-recognizer type, and (3) cluster-based, which are later combined (when multiple hits are obtained) to generate an individual's customized BRIR dataset. These can be combined, among other ways, using weighted sums. Optionally, when there is only one match, there is no need to combine the intermediate results. In one embodiment, the intermediate dataset is at least partially based on the closeness of the match of the retrieved BRIR dataset (from the candidate pool) to the extracted characteristics. In other embodiments, by using a multiple-recognizer match step, the processor retrieves one or more datasets based on multiple training parameters corresponding to the biometric data. In yet other embodiments, by using a cluster-based processing method, potential datasets are clustered based on the extracted data (e.g., biometric data). Clusters include multiple datasets that have a relationship that forms a model with corresponding BRIR datasets that match the extracted data (e.g., biometric) from the image by integral clustering or grouping.
[0031] In some embodiments of the present invention, two or more distance spheres are stored. This represents a spherical grid generated for two different distances from the listener. In one embodiment, for two or more different spherical grid distance spheres, one reference position BRIR is stored and associated. In other embodiments, each spherical grid has its own reference BRIR and will be used in combination with an applicable rotation filter. The selection processor 712 is used to match the characteristics in the memory 714 against the extracted characteristics received from the extraction device 702 for a new listener. By using various methods, the relevant characteristics are matched so that the correct BRIR dataset can be derived. As described above, these include methods for comparing biometric data based on multiple-match, multiple recognizer, and cluster-based processing methods, as well as the method described in U.S. Patent Application No. 15 / 969,767, "SYSTEM AND A PROCESSING METHOD FOR CUSTOMIZING AUDIO EXPERIENCE", filed on May 2, 2018, the entire disclosure of which is incorporated herein by reference. Column 718 represents a set of BRIR datasets of an individual measured at a second distance. That is, this column shows the BRIR dataset at the second distance recorded for the measured individual. As another example, the first BRIR dataset in column 716 can be obtained at 1.0 m to 1.5 m, while the BRIR dataset in column 718 can represent a dataset measured at 5 m from the listener. Although the BRIR datasets ideally constitute a global grid, embodiments of the present invention include subsets of BRIR pairs for conventional stereo sets, 5.1 multi-channel arrangements, 7.1 multi-channel arrangements, as well as BRIR pairs every 3° or less in both azimuth and elevation, and all other spherical grid deformations and subsets including spherical grids with irregular densities, and apply to any and all subsets of the global grid, including but not limited to all other spherical grid deformations and subsets.For example, there may be a spherical grid with a much higher density of grid points at a position in front of the listener than at a position behind the listener. Further, the configuration of the contents of columns 716 and 718 applies not only to BRIR pairs memorized from measurements and interpolation, but also to BRIR pairs further improved by generating a BRIR data set reflecting the conversion from the former to BRIRs including rotation filters.
[0032] After determining one or more matching BRIR datasets or computed BRIR datasets, these datasets are sent to the acoustic rendering device 730, and the entire BRIR dataset determined by the matching or other techniques described above for the new listener, or in some embodiments, a subset corresponding to the selected spatialized acoustic positions, is stored. Next, in one embodiment, the acoustic rendering device selects BRIR pairs at the desired azimuth or elevation positions and applies them to the input acoustic signal to provide spatial audio to the headphones 735. In other embodiments, the selected BRIR dataset is stored in a separate module coupled to the acoustic rendering device 730 and / or the headphones 735. In other embodiments, if the available capacity of the rendering device is limited, the rendering device stores only the identification information of the relevant characteristic data that best matches the listener or the identification information of the BRIR dataset that best matches, and downloads the desired BRIR pair (at the selected azimuth and elevation) in real time from the remote server 710 as needed. As described above, these BRIR pairs are derived by measurements using in-ear microphones for a medium-sized (i.e., over 100 people) population and are preferably stored together with similar image-related characteristics associated with each BRIR dataset. Instead of obtaining all 7200 points, some are generated directly by measurement and some are generated by interpolation to form a spherical grid of BRIR pairs. Even for a partially measured / partially interpolated grid, if the appropriate BRIR pairs of points from the BRIR dataset are identified using the appropriate azimuth and elevation values, other points not located on the grid lines can also be interpolated.
[0033] When a custom selected HRTF or BRIR dataset is selected for an individual, by using these personalized transfer functions, the user or system can provide at least first and second spatial acoustic positions for positioning media streams and voice communications respectively. In other words, by using a pair of transfer functions for each of the first and second spatial acoustic positions, these streams are virtually arranged, thereby enabling the listener to focus on the acoustic streams (e.g., phone or media stream) preferred by the listener due to their distinct spatial acoustic positions. The scope of the present invention is intended to cover all media streams including, but not limited to, acoustics and music associated with video.
[0034] The above invention has been described in some detail for the purpose of clear understanding, but it will be apparent that certain changes and improvements can be implemented within the scope of the appended claims. Therefore, this embodiment should be considered illustrative and not restrictive in any way. Also, the present invention is not limited to the details described herein and can be improved within the scope of the appended claims and equivalents.
Explanation of Reference Numerals
[0035] 102 First acoustic position (foreground position) 103 Headphones 104 Second acoustic position (background position) 105 Listener 110 Path 112 Intermediate point 114 Intermediate point 116 Spatial acoustic position 118 Spatial acoustic position 202 Stream 204 Stream 207 Filter 208 Filter 209 Filter 210 Filter 214 Adder 215 Adder 216 Headphone 222 Gain 223 Gain 224 Gain 225 Gain 702 Extraction Device 704 Image Sensor 706 Processor 710 Remote Server 712 Selection Processor 714 Memory 715 Column 716 Column 717 Column 718 Column 720 BRIR Generation 730 Acoustic Rendering Device 732 Memory 735 Headphone
Claims
1. An acoustic processing device that processes an event using a spatial acoustic position transfer function dataset, An acoustic rendering module configured to position a first acoustic signal and a second acoustic signal each including at least a voice communication stream and a media stream at a selected position among at least a first spatial acoustic position and a second spatial acoustic position, wherein the first spatial acoustic position and the second spatial acoustic position are each rendered using a first transfer function and a second transfer function from the spatial acoustic position transfer function dataset, respectively. An acoustic rendering module, A monitoring module that monitors the start of a voice communication event including an incoming call of a telephone call, and when the telephone call is started, positions the voice communication at the first spatial acoustic position and positions the media stream at the second spatial acoustic position, thereby processing the first acoustic signal and the second acoustic signal. And, An output module configured to render the resulting sound to headphones via two output channels. Comprising, The spatial acoustic position transfer function dataset is one of a personalized head-related impulse response (HRIR) dataset or a personalized binaural room impulse response (BRIR) dataset that is a dataset customized for an individual, An acoustic processing device that, when receiving a control signal indicating that the priority of a voice call from the individual listener is low, increases the apparent distance of the voice call and decreases the apparent distance of music using a personalized BRIR corresponding to different distances in the same direction.
2. A second processor configured to extract the individual's image-based characteristics from an input image and transmit the image-based characteristics to a selection processor configured to determine the personalized HRIR dataset or the personalized BRIR dataset from a memory having a plurality of candidate pools of HRIR or BRIR datasets provided for a group of individuals, wherein the HRIR or BRIR dataset is respectively associated with a corresponding image-based characteristic. The acoustic processing device according to claim 1, further comprising a second processor.
3. The acoustic processing device according to claim 2, wherein the selected processor determines the personalized BRIR dataset by accessing the candidate pool by comparing the extracted image-based characteristics of the individual with the extracted characteristics of the candidate pool, and identifies one or more BRIR datasets based on a measure of proximity.
4. The acoustic processing device according to claim 1, wherein the first spatial acoustic position and the second spatial acoustic position from the determined personalized BRIR dataset are derived from the captured dataset in memory by interpolation or other computational methods, and the first spatial acoustic position and the second spatial acoustic position each include a foreground position and a background position.
5. The acoustic processing device according to claim 4, wherein when receiving a control signal indicating that the priority of the voice call from the individual listener is low, the voice call is directed to the background position and the music is directed to the foreground position.
6. The acoustic processing device according to claim 1, wherein the positioning from the initial positions of the voice communication and the media stream to the first spatial acoustic position and the second spatial acoustic position is performed suddenly.
7. The acoustic processing device according to claim 1, wherein the media stream includes music.
8. The acoustic processing device according to claim 1, wherein the apparent distance of the voice call is increased and the apparent distance of the music is decreased using the first spatial acoustic position acoustic transfer function and the second spatial acoustic position acoustic transfer function from personalized BRIRs corresponding to different distances in the same direction.
9. The acoustic processing device according to claim 1, further comprising a user interface configured to select at least one location of the first spatial acoustic position and the second spatial acoustic position.
10. A method for processing an acoustic stream to headphones, comprising: Positioning at least a first acoustic signal and a second acoustic signal each including at least a voice communication stream and a media stream at a selected position of at least a first spatial acoustic position and a second spatial acoustic position, wherein the first spatial acoustic position and the second spatial acoustic position are each rendered using a first transfer function and a second transfer function from a spatial acoustic position transfer function data set, Monitoring the start of a voice communication event including an incoming telephone call, and when the telephone call is started, processing the first acoustic signal and the second acoustic signal by positioning the voice communication at the first spatial acoustic position and positioning the media stream at the second spatial acoustic position, wherein there is at least one associated room impulse response with respect to the second spatial acoustic position, Rendering the resulting sound to headphones via two output channels, comprising, wherein the spatial acoustic position transfer function data set is one of a personalized head-related impulse response (HRIR) data set or a personalized binaural room impulse response (BRIR) data set that is a data set customized for an individual, A method of increasing the apparent distance of a voice call and decreasing the apparent distance of music using a personalized BRIR corresponding to different distances in the same direction when receiving a control signal indicating that the priority of a voice call from the individual listener is low.
11. The method according to claim 10, wherein the customization includes extracting image-based characteristics of the individual from an input image and transmitting the image-based characteristics to a selection processor configured to determine a personalized HRIR data set or a personalized BRIR data set from a memory having a plurality of candidate pools of HRIR or BRIR data sets provided for a group of individuals, wherein the HRIR or BRIR data set is each associated with a corresponding image-based characteristic.
12. The method according to claim 11, wherein determining the personalized BRIR data set includes interpolation between existing BRIR data sets in the candidate pool.
Citation Information
Patent Citations
Portable terminal
JP2005269231A
Mixing techniques for mixing audio
JP2009540686A
System and method for determining head transfer function
JP2013168924A
Method for generating customized / personalized head-related transfer functions - Patent Application 20070122997
JP2019506050A