Audio processing
By detecting the obstacles input by the user, comparing the audio input of the blocked and unblocked microphone, and creating a frequency-dependent filter, it solves the problem of achieving high-quality spatial audio without using a special microphone arrangement, and realizes the effect of selectively capturing and processing audio sources.
Patent Information
- Application Number
- CN202210306872.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-26
- Filing Date
- 2022-03-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-03-25
AI Technical Summary
The prior art requires the use of special microphone arrangements when realizing high-quality spatial audio, and users hope to still obtain the advantages of spatial audio without using these special arrangements.
By detecting the presence of obstacles input by the user, receiving audio input from the blocked and unblocked microphones, comparing the two for the creation of frequency-dependent filters, and amplifying or attenuating the audio source.
It is achieved to selectively capture and process audio sources without using special microphone arrangements, providing high-quality spatial audio effects.
Smart Images

Figure CN115132216B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to audio processing. Background Art
[0002] Spatial audio supports capturing audio from an audio source while retaining information about the relative position of the audio source with respect to an origin. The audio can then be rendered to a listener at the same (or different) relative position as the listener. Specific audio sources can also be attenuated or amplified individually, or even removed completely from the rendered audio. This provides "focus" on one or more specific audio sources.
[0003] One way to capture audio while retaining information about the position of the audio source is to use a microphone array. The microphones have known fixed position differences, and thus audio from a specific audio source can arrive at each microphone with a different time delay. If the array is carefully designed, this phase information can accurately locate the audio source. Two omnidirectional microphones can locate a stationary audio source as being at the intersection (e.g., a circle) centered on the axis between the microphones. A third omnidirectional microphone locates the stationary audio source at any one of two points on either side of the plane shared by the three microphones at the intersection (circle) described above. A fourth microphone can be used to locate the stationary audio source at a single point.
[0004] Applying phase and amplitude modulation to the output of a microphone array can create a phased microphone array that can be used for beam steering of the lobes of a virtual microphone to selectively capture audio sources.
[0005] Thus it will be understood that for high-quality spatial audio, a special microphone arrangement with four or more microphones is typically used. The microphones in such an arrangement can be omnidirectional or directional.
[0006] There is a desire to obtain some of the advantages of spatial audio without the need for such a special microphone arrangement. Summary of the Invention
[0007] According to various but not necessarily all embodiments, there is provided an apparatus that includes components for:
[0008] detecting a user input indicative of the presence of at least a partially user-controlled obstacle;
[0009] receiving a blocked audio input from at least one blocked microphone when the user provides at least a partially user-controlled obstacle between the at least one blocked microphone and a first region;
[0010] receiving an unblocked audio input from at least one unblocked microphone when the user does not provide at least a partially user-controlled obstacle between the at least one unblocked microphone and the first region;
[0011] Compare the blocked audio input and the unblocked audio input;
[0012] Create a frequency - related filter based on the comparison;
[0013] Filter the audio input received from at least one microphone to create filtered audio that amplifies or attenuates the audio source in the first region.
[0014] In some but not necessarily all examples, the frequency - related filter is configured as a spatially - related filter that differentially amplifies or attenuates audio sources at different spatial positions.
[0015] In some but not necessarily all examples, the device is configured to detect a spatially - specific gesture as user input, and the spatially - specific gesture provides at least partial user control of an obstacle to the audio reaching at least one blocked microphone.
[0016] In some but not necessarily all examples, the device includes a camera and components for detecting user input indicating the presence of at least partial user control of an obstacle, and also includes components for processing the output from the camera to identify the presence of a hand and the movement of the hand as a user gesture.
[0017] In some but not necessarily all examples, the frequency - related filter is based on the spectral difference between the audio inputs between at least one blocked microphone and at least one unblocked microphone, and the spectral difference is caused by the acoustic shadow of at least partial user - controlled obstacle between at least one blocked microphone and the first region.
[0018] In some but not necessarily all examples, the device includes components for:
[0019] Spectral analysis of the blocked audio input;
[0020] Spectral analysis of the unblocked audio input;
[0021] Generate a frequency - related filter based on the difference between the spectral analyses.
[0022] In some but not necessarily all examples, the frequency - related filter selectively provides gain to spectral components corresponding to the spectral components of the blocked audio signal.
[0023] In some but not necessarily all examples, the frequency - related filter selectively provides gain to the lower - frequency harmonics of spectral components having a harmonic structure, and the lower - frequency harmonics of the spectral components are attenuated by at least partial user - controlled obstacle.
[0024] In some but not necessarily all examples, the gain is user - controllable.
[0025] In some but not necessarily all examples, the frequency - related filter is a time - varying filter configured to decay over time.
[0026] In some but not necessarily all examples, the device is configured to prompt the user to repeat the creation of the frequency - related filter, including:
[0027] Receiving blocked audio input from at least one blocked microphone when the user provides at least a partial obstacle between the at least one blocked microphone and the first region;
[0028] Receiving unblocked audio input from at least one unblocked microphone when the user does not provide at least a partial obstacle between the at least one unblocked microphone and the first region;
[0029] Comparing the blocked audio input and the unblocked audio input;
[0030] Creating a frequency - related filter based on the comparison;
[0031] Filtering the audio input received from at least one microphone to create filtered audio that amplifies or attenuates an audio source in the first region.
[0032] According to various but not necessarily all embodiments, a method is provided that includes:
[0033] Receiving blocked audio input from at least one blocked microphone when the user provides at least a partial user - controlled obstacle between the at least one microphone and the first region;
[0034] Receiving unblocked audio input from at least one unblocked microphone when the user does not provide at least a partial user - controlled obstacle between the at least one microphone and the first region;
[0035] Comparing the blocked audio input and the unblocked audio input;
[0036] Creating a frequency - related filter based on the comparison;
[0037] Filtering the audio input received from at least one microphone to create filtered audio that amplifies or attenuates an audio source in the first region.
[0038] In some but not necessarily all examples, the frequency - related filter is based on the spectral difference between the audio inputs from the blocked microphone and the unblocked microphone, which is caused by the acoustic shadow of the user's hand.
[0039] In some but not necessarily all examples, the frequency - related filter provides a frequency - related gain, and the gain is controlled by the user to decay or amplify.
[0040] According to various but not necessarily all embodiments, a computer program is provided that, when run on at least one processor, performs:
[0041] Detecting a user input indicating the presence of at least a partially user-controlled obstacle;
[0042] Receiving blocked audio input from at least one blocked microphone when the user provides at least a partially user-controlled obstacle between the at least one blocked microphone and a first region;
[0043] Receiving unblocked audio input from at least one unblocked microphone when the user does not provide at least a partially user-controlled obstacle between the at least one unblocked microphone and the first region;
[0044] Comparing the blocked audio input and the unblocked audio input;
[0045] Creating a frequency-dependent filter based on the comparison;
[0046] Filtering the audio input received from at least one microphone to create filtered audio that amplifies or attenuates an audio source in the first region.
[0047] According to various but not necessarily all embodiments, a device is provided that includes components for:
[0048] Detecting a user input indicating the presence of at least a partially user-controlled obstacle;
[0049] Receiving blocked audio input from at least one microphone when the user provides at least a partially user-controlled obstacle;
[0050] Receiving unblocked audio input from at least one microphone when the user does not provide at least a partially user-controlled obstacle;
[0051] Comparing the blocked audio input and the unblocked audio input;
[0052] Creating a frequency-dependent filter based on the comparison;
[0053] Filtering the audio input received from at least one microphone to create filtered audio that amplifies or attenuates an audio source when the user does not provide at least a partially user-controlled obstacle.
[0054] In one embodiment, a device includes components for:
[0055] Detecting a user input indicating the presence of at least a partially user-controlled obstacle;
[0056] Receiving blocked audio input from a microphone when the user provides at least a partially user-controlled obstacle;
[0057] When the user does not provide at least a partial user-controlled obstruction, receive unobstructed audio input from a microphone;
[0058] Compare the obstructed audio input and the unobstructed audio input;
[0059] Create a frequency-dependent filter based on the comparison;
[0060] Filter the audio input received from the microphone to create filtered audio.
[0061] In another embodiment, an apparatus includes components for:
[0062] Detect a user input indicating the presence of at least a partial user-controlled obstruction;
[0063] When the user provides at least a partial user-controlled obstruction, receive obstructed audio input from a first microphone;
[0064] Receive unobstructed audio input from a second microphone;
[0065] Compare the obstructed audio input and the unobstructed audio input;
[0066] Create a frequency-dependent filter based on the comparison;
[0067] Filter the audio input received from the first microphone and / or the second microphone to create filtered audio.
[0068] According to various but not necessarily all embodiments, examples are provided as claimed in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Some examples will now be described with reference to the drawings, in which:
[0070] Figure 1A and Figure 1B show examples of obstructing an audio source to create a frequency-dependent filter;
[0071] Figure 2A and Figure 2B show examples of obstructing an audio source to create a frequency-dependent filter;
[0072] Figure 2C show examples of obstructing an audio source to create a frequency-dependent filter;
[0073] Figure 3 show examples of methods of creating a frequency-dependent filter;
[0074] Figure 4 show examples of methods of creating a frequency-dependent filter after a user input;
[0075] Figure 5A , Figure 5B and Figure 5C respectively show examples of the spectrum of an unobstructed audio input, an example of the spectrum of an obstructed audio input, and an example of the difference between the spectrum of an unobstructed audio input and the spectrum of an obstructed audio input;
[0076] Figure 6A shows an example of an amplification filter based on Figure 5C the difference between the spectrum of the unobstructed audio input and the spectrum of the obstructed audio input shown;
[0077] Figure 6B shows an example of an attenuation filter based on Figure 5C the difference between the spectrum of the unobstructed audio input and the spectrum of the obstructed audio input shown;
[0078] Figure 7 shows an example of an apparatus for creating a frequency-dependent filter;
[0079] Figure 8 shows an example of an apparatus for using a frequency-dependent filter;
[0080] Figure 9A shows an example of an implementation of an apparatus;
[0081] Figure 9B shows an example of a computer program capable of creating a frequency-dependent filter;
[0082] Figure 10A , Figure 10B and Figure 10C respectively show examples of the spectrum of an unobstructed audio input, an example of the spectrum of an obstructed audio input having a harmonic structure, and an example of the difference between the spectrum of an unobstructed audio input and the spectrum of an obstructed audio input;
[0083] Figure 11A shows an example of a harmonic expansion amplification filter based on Figure 10C the difference between the spectrum of the unobstructed audio input and the spectrum of the obstructed audio input shown; and
[0084] Figure 11B shows an example of a harmonic expansion attenuation filter based on Figure 10C the difference between the spectrum of the unobstructed audio input and the spectrum of the obstructed audio input shown. DETAILED DESCRIPTION
[0085] In this specification, reference numerals without subscripts may be used to identify a group of objects or features. Particular members of the group (if more than one) may (but need not) be identified using reference numerals with subscripts.
[0086] In this specification, the reference numeral 22' will be used to refer to user-impeded audio input to distinguish it from other audio inputs 22. The reference numeral 22 will be used to refer to audio input in the absence of a user obstacle 30. The audio input in the absence of a user obstacle 30 can be used for comparison purposes to create a filter 50 (describing the audio input 22 as an "unimpeded" audio input at this stage to distinguish it from the impeded audio input 22'). The audio input in the absence of an intentional user obstacle 30 can be used for the purpose of being filtered by the created filter 50.
[0087] The following description describes various examples in which the apparatus 100 includes components for:
[0088] detecting a user input 134 indicative of the presence of a user-controlled obstacle 30;
[0089] receiving an impeded audio input 22' from at least one impeded microphone 20 when the user 130 provides at least a partial obstacle 30 between the at least one impeded microphone 20 and the first region 16;
[0090] receiving an unimpeded audio input 22 from at least one unimpeded microphone 20 when the user 130 does not provide at least a partial obstacle 30 between the at least one unimpeded microphone 20 and the first region 16;
[0091] comparing the impeded audio input 22' and the unimpeded audio input 22;
[0092] creating a frequency-dependent filter 50 based on the comparison; and
[0093] filtering the audio input 22 received from at least one microphone 20 to create a filtered audio that amplifies or attenuates an audio source 10 in the first region 16.
[0094] In some but not necessarily all examples, the frequency-dependent filter 50 is based on the spectral difference of the audio input 22 between at least one unimpeded microphone and at least one impeded microphone 20, which is caused by the acoustic shadow of at least a partial obstacle 30 between the at least one impeded microphone 20 and the first region 16. In some examples, at least a partial obstacle 30 is the hand 132 of the user 130.
[0095] In some but not necessarily all examples, the frequency-dependent filter 50 provides a frequency-dependent gain. The gain can be a negative gain (attenuation) or a positive gain (amplification). In some examples, the gain is controlled by the user 130 to attenuate or amplify. In some examples, the user provides this control via a gesture 134 of their hand 132.
[0096] Figure 1A 、Figure 1B , Figure 2A , Figure 2B , Figure 2C The corresponding audio 12 is shown i One or more spatially distributed audio sources 10 i The audio source 10 shown in the figure i and their arrangement are only examples. There may be more or fewer audio sources 10, for example there may be a single audio source 10. The location of the audio source 10 may be different from that shown. In addition, the audio source 10 may be positioned in a three-dimensional space. The audio source 10 may have a different spatial range than that shown. In some examples, the audio source may be an ambient audio source (e.g., background noise). The audio source 10 may be a localized audio source (e.g., human speech).
[0097] exist Figure 1A , Figure 1B In FIG. 5 , a frequency dependent filter 50 is created based on the audio input 22 from a single microphone 20 (not shown in these figures). Figure 1B ) and unobstructed audio input 22 ( Figure 1A ) is based on audio inputs 22, 22' captured by microphone 20 at different times. The comparison is a time-divided comparison.
[0098] exist Figure 2A , Figure 2B , Figure 2C In FIG. 5 , a frequency dependent filter 50 is created based on audio input 22 from two microphones 201 , 202 (not shown in these figures), which are spatially distinct and separated. Figure 2B ) and unobstructed audio input 22 ( Figure 2A ) can be based on the audio inputs 22, 22' captured by the microphone 20 at different times, such as Figure 1A , Figure 1B However, the blocked audio input 222' ( Figure 2C ) and unobstructed audio input 22 ( Figure 2C ) can be based on audio inputs 221, 222' captured by different microphones 201, 202 at the same time (eg, simultaneously or concurrently). The comparison is space-divided because the different microphones 201, 202 are at different locations.
[0099] exist Figure 1A , the microphones 20 provide unobstructed audio input 22 captured at times (offset times) when a user does not provide at least a partial obstruction 30 between the at least one microphone 20 and the first area 16 for further processing.
[0100] In Figure 1B , the user provides an obstacle 30 between at least one microphone 20 and a first region 16. The obstacle 30 can be a complete or partial obstacle. The obstacle 30 obstructs the audio 122 generated by the audio source 102. The audio that reaches the microphone 20 (if any) from the audio source 102 is the obstructed audio 122'. At the time (reference time) when the user provides at least a partial obstacle 30 between the microphone 20 and the first region 16, the microphone 20 captures the obstructed audio 122' and the unobstructed audio 121, 123 from other audio sources 101, 103 (if any) to generate an obstructed audio input 22'. The microphone 20 provides the obstructed audio input 22' for further processing.
[0101] As described later, a device 100, which may or may not include the microphone 20, compares the obstructed audio input 22' and the unobstructed audio input 22, and creates a frequency-dependent filter 50 based on the comparison. After the user provides that the obstacle 30 has been removed, the created frequency-dependent filter 50 can then be used to filter the audio input 22 received from the microphone 20 to create a filtered audio that amplifies or attenuates the audio source in the first region 16.
[0102] The offset time is a time offset relative to the reference time. The offset time can be before or after the reference time. The offset time can be immediately before or after the reference time.
[0103] Thus, at the reference time, the user creates a spatially related at least partial obstacle 30 to the audio reaching at least one microphone. The user input indicating the reference time can be used to indicate the user's control of the presence of the obstacle 30 at the reference time. The device 100 can be configured to detect the user input determining the reference time, compare the audio input 22 received from the microphone 20 at the reference time with the audio input 22 received from the microphone 20 at an offset time (a time offset relative to the reference time), and create a frequency-dependent filter 50 based on the comparison.
[0104] At the offset time, the microphone 20 is an unobstructed microphone 20.
[0105] At the reference time, the microphone 20 is an obstructed microphone 20.
[0106] In Figure 2A , the microphone 20 provides unobstructed audio inputs 221, 222 captured at a time (offset time) when the user does not provide at least a partial obstacle 30 between at least one microphone 20 and the first region 16 for further processing. In this example, a pair of microphones 201, 202 is used, but more microphones can be used.
[0107] In Figure 2BIn this case, the user provides an obstacle 30 between a pair of microphones 20 and the first region 16. The obstacle 30 can be a complete or partial obstacle. The obstacle 30 obstructs the audio 122 generated by the audio source 102.
[0108] The audio from the audio source 102 that reaches the microphone 201 (if any) is the obstructed audio 122'. At the time (reference time) when the user provides at least a partial obstacle 30 between the microphone 201 and the first region 16, the microphone 201 captures the obstructed audio 122' and the unobstructed audio 121, 123 from other audio sources 101, 103 (if any) to generate an obstructed audio input 221'. The microphone 201 provides the obstructed audio input 221' for further processing.
[0109] The audio from the audio source 102 that reaches the microphone 202 (if any) is the obstructed audio 122'. At the time (reference time) when the user provides at least a partial obstacle 30 between the microphone 202 and the first region 16, the microphone 202 captures the obstructed audio 122' and the unobstructed audio 121, 123 from other audio sources 101, 103 (if any) to generate an obstructed audio input 222'. The microphone 202 provides the obstructed audio input 222' for further processing.
[0110] As described later, the device 100 that may or may not include the (multiple) microphones 20 compares the obstructed audio inputs 22' (the obstructed audio input 221' and the obstructed audio input 222') and the unobstructed audio inputs 22 (the unobstructed audio input 221 and the unobstructed audio input 222), and creates a frequency - related filter 50 based on the comparison. In some examples, the combination (e.g., sum) of the obstructed audio input 221' and the obstructed audio input 222' is compared with the combination (e.g., sum) of the unobstructed audio input 221 and the unobstructed audio input 222.
[0111] After the obstacle 30 provided by the user has been removed, the created frequency - related filter 50 can then be used to filter the audio input 22 received from the microphone 20 to create a filtered audio that amplifies or attenuates the audio source 102 in the first region 16.
[0112] The offset time is the time offset relative to the reference time. The offset time can be before or after the reference time. The offset time can be immediately before or after the reference time.
[0113] Thus, at a reference time, the user creates at least a partially space - related obstacle 30 to the audio arriving at the microphone 20. A user input indicating the reference time can be used to indicate user control of the presence of the obstacle 30 at the reference time. The apparatus 100 can be configured to detect the user input that determines the reference time, compare the audio input 22 received from the microphone 20 at the reference time with the audio input 22 received from the microphone 20 at an offset time (which has a time offset relative to the reference time), and create a frequency - related filter 50 based on the comparison.
[0114] At the offset time, the microphones 201, 202 are unobstructed microphones 20.
[0115] At the reference time, the microphones 201, 202 are obstructed microphones 20
[0116] At Figure 2C , the microphone 20 provides an obstructed audio input and an unobstructed audio input for further processing. In this example, a pair of microphones 201, 202 is used, but more microphones can be used. The user provides the obstacle 30 between the microphone 202 and the first region 16 (but not between the microphone 201 and the first region 16). The obstacle 30 can be a complete or partial obstacle. The obstacle 30 obstructs the audio 122 generated by the audio source 102 and captured by the microphone 202.
[0117] When the user provides the obstacle 30, the microphone 201 provides an unobstructed audio input 221 for further processing, and the microphone 202 provides an obstructed audio input 222' for further processing.
[0118] In this example, the user controls the obstacle 30 to be between the microphone 202 and the first region 16, but not between the microphone 201 and the first region 16. Thus, the microphone 201 provides the unobstructed audio input 221 captured when the user does not provide at least a partial obstacle 30 between the microphone 201 and the first region 16 for further processing, and the microphone 202 provides the obstructed audio input 222' captured when the user provides at least a partial obstacle 30 between the microphone 202 and the first region 16 for further processing.
[0119] The audio arriving at the microphone 201 (if any) from the audio source 102 is unobstructed audio 122. At the time (reference time) when the user provides at least a partial obstacle 30 between the microphone 202 and the first region 16, the microphone 201 captures unobstructed audio 122 and unobstructed audio 121, 123 from other audio sources 101, 103 (if any) to generate the unobstructed audio input 221. The microphone 201 provides the unobstructed audio input 221 for further processing.
[0120] The audio that reaches microphone 202 (if any) from audio source 102 is blocked audio 122'. At the time (reference time) when the user provides at least a partial obstacle 30 between microphone 202 and the first region 16, microphone 202 captures blocked audio 122' and unblocked audio 121, 123 from other audio sources 101, 103 (if any) to generate a blocked audio input 222'. Microphone 202 provides the blocked audio input 222' for further processing.
[0121] As described later, apparatus 100, which may or may not include some or all of microphones 20, compares the blocked audio input 22' (blocked audio input 222') and the unblocked audio input 22 (unblocked audio input 221), and creates a frequency - related filter 50 based on the comparison.
[0122] After the user provides that the obstacle 30 has been removed, the created frequency - related filter 50 can then be used to filter the audio input 22 received from (a plurality of) microphones 20 to create filtered audio that amplifies or attenuates the audio sources in the first region 16.
[0123] Thus, at the reference time, the user creates at least a partial, spatially - related obstacle 30 to the audio reaching microphone 202. The user input indicating the reference time can be used to indicate the user's control of the presence of obstacle 30 at the reference time. Apparatus 100 can be configured to detect the user input that determines the reference time, compare the audio input 222 received from microphone 202 at the reference time with the audio input 221 received from microphone 201 at the reference time, and create a frequency - related filter 50 based on the comparison.
[0124] At the reference time, microphone 202 is a blocked microphone 20, and microphone 201 is an unblocked microphone 20
[0125] Figure 3 Method 200 is shown. Method 200 creates a frequency - related filter 50 (not shown).
[0126] In block 210, method 200 includes receiving a blocked audio input 22' from at least one blocked microphone 20 when the user provides at least a partial obstacle 30 between the at least one blocked microphone 20 and the first region 16.
[0127] In block 212, method 200 includes receiving an unblocked audio input 22 from at least one unblocked microphone 20 when the user does not provide at least a partial obstacle 30 between the at least one unblocked microphone 20 and the first region 16.
[0128] In block 214, method 200 includes comparing the blocked audio input 22' and the unblocked audio input 22.
[0129] At block 216, method 200 includes creating a frequency - related filter 50 based on a comparison.
[0130] At block 218, method 200 includes filtering an audio input received from at least one microphone 20 to create filtered audio that amplifies or attenuates an audio source in the first region 16.
[0131] The filtering at block 218 occurs after the removal of the obstacle 30 mentioned at block 210.
[0132] The at least one blocked microphone 20 and the at least one unblocked microphone 20 can be the same set of one or more microphones 20 at different times (time - division).
[0133] The at least one blocked microphone 20 and the at least one unblocked microphone 20 can be different sets of one or more microphones 20 at the same time (space - division).
[0134] For example, the frequency - related filter 50 can be based on a spectral difference of the audio input 22 from at least one microphone 20, which is caused by the acoustic shadow of the user's hand. The difference can be a change over time of one or more microphones 20 (time - division). The difference can be a difference between microphones 20 at the same time (space - division).
[0135] Figure 4 A method 201 similar to method 200 is shown. Method 201 creates a frequency - related filter 50 (not shown).
[0136] At block 202, method 201 includes detecting a user input that determines a reference time at which the user provides at least a partial obstacle 30 between at least one blocked microphone 20 and the first region 16, but does not provide an obstacle 30 between at least one unblocked microphone 20 and the first region 16.
[0137] At block 210, method 201 includes receiving a blocked audio input 22' from at least one blocked microphone 20 when the user provides at least a partial obstacle 30 between the at least one blocked microphone 20 and the first region 16.
[0138] At block 212, method 201 includes receiving an unblocked audio input 22 from at least one unblocked microphone 20 when the user does not provide at least a partial obstacle 30 between the at least one unblocked microphone 20 and the first region 16.
[0139] At block 214, method 201 includes comparing the blocked audio input 22' and the unblocked audio input 22.
[0140] At block 216, method 201 includes creating a frequency-dependent filter 50 based on a comparison.
[0141] At block 218, method 201 includes filtering an audio input received from at least one microphone 20 to create filtered audio that amplifies or attenuates an audio source in the first region 16.
[0142] The filtering at block 218 occurs after the removal of the obstacle 30 mentioned at block 210.
[0143] The at least one blocked microphone 20 and the at least one unblocked microphone 20 can be the same set of one or more microphones 20 at different times (time division). In this example, at block 210, method 201 includes receiving a blocked audio input 22' captured by the (multiple) microphones 20 at a reference time when the user provides at least a partial obstacle 30 between the at least one blocked microphone 20 and the first region 16, and at block 212, method 201 includes receiving an unblocked audio input 22 captured by the (multiple) microphones 20 at an offset time (a time offset relative to the reference time) when the user does not provide at least a partial obstacle 30 between the at least one unblocked microphone 20 and the first region 16.
[0144] The at least one blocked microphone 20 and the at least one unblocked microphone 20 can be different sets of one or more microphones 20 at the same time (space division). In this example, at block 210, method 201 includes receiving a blocked audio input 22' captured by a first set of one or more microphones 20 at a reference time when the user provides at least a partial obstacle 30 between the first set of microphones 20 and the first region 16, and at block 212, method 201 includes receiving an unblocked audio input 22 captured by a second set of one or more microphones 20 at the reference time when the user does not provide at least a partial obstacle 30 between the second set of microphones 20 and the first region 16.
[0145] Figure 5A is an example of a spectral representation of the unblocked audio input 22 produced by the (multiple) microphones 20 that capture the unblocked audio 12. For purposes of explanation, this example is simplified. The figure shows the energy (y-axis) that the unblocked audio input 22 has within a frequency band (x-axis). Although the frequency bands have the same size, in other examples, the frequency bands can have different sizes. There can also be more or fewer frequency bands.
[0146] Thus, Figure 5A represents the spectrum of the unblocked audio input 22.
[0147] Figure 5Bis an example of a spectral representation of a blocked audio input 22' generated by one or more microphones 20 that capture blocked audio 12', where the blocked audio 12' is a blocked variant of unblocked audio 12. Thus, Figure 5B represents the spectrum of the blocked audio input 22'.
[0148] Figure 5C shows the spectrum of the unblocked audio input 22 ( Figure 5A ) and the difference 40 between the spectrum of the blocked audio input 22' ( Figure 5B ). In this example, but not necessarily in all examples, the spectrum of the blocked audio input 22' ( Figure 5A ) is subtracted from the spectrum of the unblocked audio input 22 ( Figure 5B ). In this example, the difference 40 represents the spectrum of the audio that has been blocked from the first region 16. The spectrum of the audio that has been blocked from the first region 16 has a range 42.
[0149] For example, the comparison described above can determine the difference 40 and use it to create a filter 50.
[0150] The spectrum of the audio that has been blocked from the first region 16 (e.g., Figure 5C ) is transformed into Figure 6A and Figure 6B in a frequency filter 50. The filter applies a gain to the spectral range 42.
[0151] In Figure 6A , the gain is a positive gain (>1) within the range 42 and a negative gain (<1) outside the range 42. Thus, Figure 6A shows an amplification filter 50 that is configured to preferentially amplify the spectrum of the audio that has been blocked from the first region 16.
[0152] In Figure 6B , the gain is a negative gain (<1) within the range 42 and a positive gain (>1) outside the range 42. Thus, Figure 6B shows an attenuation filter 50 that is configured to preferentially attenuate the spectrum of the audio that has been blocked from the first region 16.
[0153] Thus, it should be understood that the frequency-dependent filter 50 is configured as a spatially-dependent filter 50 that differentially amplifies or attenuates audio sources 10 at different spatial locations.
[0154] The frequency - dependent filter 50 is based on the spectral difference 40 between the blocked audio input 22' caused by the acoustic shadow of the obstacle 30 and the unblocked audio input 22. The obstacle 30 can be, for example, the hand 132 of the user 130, and the hand 132 provides at least a partial obstacle 30 between the blocked microphone 20 and the first region 16.
[0155] The frequency - dependent filter 50 selectively provides gain for spectral components in the range 42 that correspond to the spectral components of the audio signal that has been blocked and not captured. The gain can be controlled by the user 130. For example, in some but not necessarily all examples, the user 130 can select whether the filter is an amplification filter or an attenuation filter. For example, in some but not necessarily all examples, the user 130 can determine the magnitude of the gain, e.g., the magnitude of the amplification / attenuation. In some examples, the user input can be via a gesture 134 of the hand 132 (see Figure 7 ). The gesture can be part of or separate from the gesture used to place the hand 132 as the obstacle 30.
[0156] In this example, the frequency - dependent filter 50 can be based, for example, on the spectral difference of the audio input 22 from at least one microphone 20, which is caused by the acoustic shadow of the hand 132 of the user 130. The difference 40 can be a change over time (time - division) of one or more microphones 20. The difference 40 can be a difference between microphones 20 at the same time (space - division).
[0157] Figure 7 An example of the apparatus 100 is shown, and the apparatus 100 includes:
[0158] A component 112 for detecting a user input 134 indicating the presence of a user - controlled obstacle 30;
[0159] Components 102, 104 for receiving a blocked audio input 22' from at least one blocked microphone 20 when the user 130 provides at least a partial obstacle 30 between the at least one blocked microphone 20 and the first region 16;
[0160] Components 102, 104 for receiving an unblocked audio input 22 from at least one unblocked microphone 20 when the user 130 does not provide at least a partial obstacle 30 between the at least one unblocked microphone 20 and the first region 16;
[0161] A component 104 for comparing the blocked audio input 22' and the unblocked audio input 22; and
[0162] A component 106 for creating a frequency - dependent filter 50 based on the comparison (e.g., based on the difference 40).
[0163] As Figure 8As shown, the apparatus 100 may also include components (frequency - dependent filter 50) for filtering the audio input 22 received from at least one microphone 20 to create filtered audio 24 that amplifies or attenuates the audio source in the first region 16.
[0164] In this example, the apparatus 100 additionally includes a spectral analysis component 102 that performs spectral analysis on the blocked audio input 22' received from the blocked microphone(s) 20 and spectral analysis on the unblocked audio input 22 received from the unblocked microphone(s) 20. The spectral analysis component can be, for example, a spectrum analyzer configured to create an unblocked spectrum of the unblocked audio input 22 by converting the received unblocked audio input 22 from the time domain to the frequency domain (e.g., as Figure 5A shown), and create a blocked spectrum of the blocked audio input 22' by converting the received blocked audio input 22' from the time domain to the frequency domain (e.g., as Figure 5B shown). At block 106, a frequency - dependent filter 50 is generated based on the difference 40 between the blocked spectrum and the unblocked spectrum.
[0165] In some time - division examples, the apparatus 100 is configured to assume an obstacle 30 at a reference time indicated by the user and assume no obstacle 30 at an offset time. The blocked spectrum is obtained by analyzing the audio input from the microphone(s) 20 captured at the reference time, and the unblocked spectrum is obtained by analyzing the audio input from those microphone(s) 20 captured at the offset time.
[0166] The control block 110 can be configured to control when to use or not use the described method.
[0167] The control block 110 can be configured to identify user input.
[0168] A sensor 112 can be provided to detect user input. For example, the sensor 112 (not the microphone 20) can detect the movement 134 of the hand 132 of the user 130. In some examples, the sensor 112 can position the hand 132 of the user 130 relative to the microphone 20. Any suitable sensor 112 can be used.
[0169] In some examples, the control block 110 is configured to identify that the obstacle 30 is the hand 132 of the user 130. Differential hand masking results in a frequency difference in the audio spectrum. In some examples, the gain is controlled by the user 130 to attenuate or amplify. In some examples, the user provides this control via the gesture 134 of their hand 132.
[0170] For example, user 130 can control the preferred behavior for obstruction by performing an assistive gesture 134 with the obstructing hand 132. An example can be first placing the hand as an obstacle 30, and once the device 100 provides an output confirming the action, the user can perform a zoom gesture (e.g., an open gesture performed by moving the thumb and index finger away from each other) or an attenuation gesture (e.g., a pinch gesture performed by moving the thumb and index finger closer to each other) to control whether the filter 50 should be a zoom filter or an attenuation filter, respectively. The degree of zoom / attenuation can be controlled by the size of the zoom gesture (e.g., the amount by which the distance between the thumb and index finger increases) or the size of the attenuation gesture (e.g., the amount by which the distance between the thumb and index finger decreases). The device 100 can, for example, provide feedback to the user 130 regarding the zoom / attenuation gesture.
[0171] In some examples, the sensor 112 is a camera. Examples of other sensors 112 can be proximity sensors, hover sensors, positioning sensors, or some other sensor.
[0172] The camera 112 can, for example, be a camera that records the visual part of an audiovisual scene. The microphone 20 can simultaneously record the audio part of the audiovisual scene. As described above, the audio part can be filtered by the frequency-dependent filter 50.
[0173] The sensor 112 can be configured to detect a spatially specific gesture 134 as user input, where the spatially specific gesture 134 provides at least a partially spatially related obstacle 30 to the audio reaching at least one microphone.
[0174] For example, in some examples, the control block 110 can be configured to process the output from the camera 112 to recognize the presence of the hand 132, the position of the hand 132, and the movement of the hand 132 as user gestures 134. For example, the control block 110 can distinguish different gestures 134 as different user inputs. Computer vision processing can be used to distinguish the hand 132 from other objects and to distinguish different gestures 134.
[0175] The user 130 can indicate the preferred audio focus direction by placing his / her hand between the microphone 20 and the preferred audio focus direction. Even during audio / video capture, the gesture 134 provides an intuitive way to select the focus position (i.e., select the first region 16). The focus direction selection can be performed without manual interaction with the device 100, which is particularly useful, for example, when wearing winter gloves.
[0176] The situation that leads to the generation of the specific frequency-dependent filter 50 can change over time. If the frequency-dependent filter 50 is not changed, replaced, or updated, it can produce incorrect or undesirable results in some cases.
[0177] In one example, the device 100 continues to perform spectral analysis on the audio input 22 to detect changes in the configuration of the audio source 10, such as changes, additions, losses, or movements of the audio source 10. The spectral analysis can be limited, for example, to the spectral range 42 of the unobstructed audio input 22 to detect substantial changes in that spectral region. The device can then prompt the user to perform a recalibration process. The recalibration process is, for example, a repetition of the original process for creating the frequency-dependent filter 50.
[0178] In this example or other examples, the device 100 can change the frequency-dependent filter over time or automatically stop using the frequency-dependent filter 50. For example, the frequency-dependent filter 50 can be a time-variable filter 50 configured to decay over time. For example, the magnitude of the gain difference provided by the filter 50 can be time-dependent and decrease to zero over time (e.g., 10 seconds). The decrease of the filter 50 can be visually shown by the device 100, for example. If still needed, the user then has to re-perform the process of creating the frequency-dependent filter 50.
[0179] Thus, the device 100 can be configured to prompt the user to perform "recalibration" by repeating the process of creating the frequency-dependent filter 50. This can include, for example:
[0180] Receiving an obstructed audio input 22' from at least one microphone 20 when the user provides at least a partial obstruction 30 between at least one microphone 20 and a region;
[0181] Receiving an unobstructed audio input 22 from at least one microphone 20 when the user does not provide at least a partial obstruction 30 between at least one microphone 20 and a first region 16;
[0182] Comparing the obstructed audio input and the unobstructed audio input;
[0183] Creating the frequency-dependent filter 50 based on the comparison;
[0184] Filtering the audio input 22 received from at least one microphone 20 to create filtered audio that amplifies or attenuates the audio source in the first region 16.
[0185] In some of the described examples but not necessarily all of the described examples, the device 100 is a handheld portable device that can fit in a jacket pocket. In some of the described examples but not necessarily all of the described examples, the device 100 is a flat-screen tablet device, such as a mobile phone, a tablet computer, a personal digital assistant, etc.
[0186] Figure 8An example of apparatus 100 is shown that uses a frequency-dependent filter 50 to filter an audio input 22 received from a (plurality of) microphones 20 to create a filtered audio 24 that amplifies or attenuates an audio source in a first region 16.
[0187] The same frequency-dependent filter 50 can be used to filter all audio inputs 22 received from the microphones 20 to create a filtered audio 24 that amplifies or attenuates an audio source in the first region 16. Alternatively, the same frequency-dependent filter 50 can be used to filter a subset of the audio inputs 22 received from the microphones 20 to create a filtered audio 24 that amplifies or attenuates an audio source in the first region 16.
[0188] Figure 8 It is shown that the received audio input 22 can be preprocessed at a preprocessing block 120 before being filtered. In this example, the received audio input 22 can also be preprocessed before performing a spectral analysis 102 ( Figure 7 ).
[0189] For example, the preprocessing can include noise reduction, equalization, spatialization of the microphone signals, wind noise reduction, etc.
[0190] The filtered audio input 22 can be “live,” i.e., real-time, or can be accessed from a memory.
[0191] The audio output (filtered audio 24) can be rendered “live,” i.e., real-time, or can be recorded in a memory for future access.
[0192] Figure 9A An example of a controller 70 for the apparatus 100 is shown. The implementation of the controller 70 can be as a controller circuitry. The controller 70 can be implemented solely in hardware, have certain aspects implemented in software including only firmware, or can be a combination of hardware and software (including firmware).
[0193] As Figure 9A shown, the controller 70 can be implemented using instructions that enable hardware functions, e.g., by using executable instructions of a computer program 76 in a general-purpose or a special-purpose processor 72, which can be stored on a computer-readable storage medium (disk, memory, etc.) for execution by such a processor 72.
[0194] The processor 72 is configured to read from and write to a memory 74. The processor 72 can also include an output interface through which the processor 72 outputs data and / or commands, and an input interface through which data and / or commands are input to the processor 72.
[0195] The memory 74 stores a computer program 76 including computer program instructions (computer program code) which, when loaded into the processor 72, control the operation of the apparatus 100. The computer program instructions of the computer program 76 provide the logic and routines that enable the apparatus to execute the method shown in the figures. The processor 72 is capable of loading and executing the computer program 76 by reading the memory 74.
[0196] Accordingly, the apparatus 100 comprises:
[0197] at least one processor 72; and
[0198] at least one memory 74 including computer program code,
[0199] the at least one memory 74 and the computer program code are configured to, together with the at least one processor 72, cause the apparatus 100 to at least perform:
[0200] detect a user input indicating the presence of a user-controlled obstacle 30;
[0201] receive a blocked audio input 22' from the at least one microphone 20 when the user provides at least a partial obstacle 30 between the at least one microphone 20 and the first region 16;
[0202] receive an unblocked audio input 22 from the at least one microphone 20 when the user does not provide at least a partial obstacle 30 between the at least one microphone 20 and the first region 16;
[0203] compare the blocked audio input and the unblocked audio input;
[0204] create a frequency-dependent filter 50 based on the comparison;
[0205] filter the audio input 22 received from the at least one microphone 20 to create filtered audio that amplifies or attenuates an audio source in the first region 16.
[0206] As Figure 9B shown, the computer program 76 can reach the apparatus 100 via any suitable delivery mechanism 78. The delivery mechanism 78 can be, for example, a machine-readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a memory device, a recording medium (such as a compact disc read-only memory (CD-ROM) or a digital versatile disc (DVD) or a solid-state memory), an article of manufacture including or tangibly embodying the computer program 76. The delivery mechanism can be a signal configured to reliably convey the computer program 76. The apparatus 100 can propagate or transmit the computer program 76 as a computer data signal.
[0207] Computer program instructions for causing an apparatus to perform at least the following operations or for performing at least the following operations:
[0208] Detect a user input indicating the presence of a user-controlled obstacle 30;
[0209] Receive a blocked audio input 22' from at least one microphone 20 when the user provides at least a partial obstacle 30 between the at least one microphone 20 and a first region 16;
[0210] Receive an unblocked audio input 22 from at least one microphone 20 when the user does not provide at least a partial obstacle 30 between the at least one microphone 20 and the first region 16;
[0211] Compare the blocked audio input and the unblocked audio input;
[0212] Create a frequency-dependent filter 50 based on the comparison;
[0213] Filter an audio input 22 received from at least one microphone 20 to create a filtered sound that amplifies or attenuates an audio source in the first region 16.
[0214] The computer program instructions may be included in a computer program, a non-transitory computer-readable medium, a computer program product, a machine-readable medium. In some but not necessarily all examples, the computer program instructions may be distributed over more than one computer program.
[0215] Although the memory 74 is shown as a single component / circuit system, it may be implemented as one or more separate component / circuit systems, some or all of which may be integrated / removable and / or may provide permanent / semi-permanent / dynamic / cache storage.
[0216] Although the processor 72 is shown as a single component / circuit system, it may be implemented as one or more separate component / circuit systems, some or all of which may be integrated / removable. The processor 72 may be a single-core or multi-core processor.
[0217] Figure 10A 、 Figure 10B 、 Figure 10C 、 Figure 11A 、 Figure 11B Corresponding to the previous Figure 5A 、 Figure 5B 、 Figure 5C 、 Figure 6A 、 Figure 6B 。
[0218] Figure 10A 、 Figure 10B 、 Figure 10C 、Figure 11A shows that the frequency-dependent filter 50 ( Figure 6A shown) can be extended to apply amplification outside of range 42 at lower frequency harmonic frequencies. Figure 10A , Figure 10B , Figure 10C , Figure 11B shows that the frequency-dependent filter 50 ( Figure 6B shown) can be extended to apply attenuation outside of range 42 at lower frequency harmonic frequencies.
[0219] Figure 10A is an example of an unobstructed spectrum, which is the spectrum of the unobstructed audio input 22 generated by the (one or more) microphones 20 that capture the unobstructed audio 12. The unobstructed spectrum includes a harmonic structure (H).
[0220] Figure 10B is an example of an obstructed spectrum, which is the spectrum of the obstructed audio input 22' generated by the (one or more) microphones 20 that capture the obstructed audio 12'.
[0221] Figure 10C shows the difference 40 between the unobstructed spectrum of the unobstructed audio input 22 ( Figure 10A ) and the obstructed spectrum of the obstructed audio input 22' ( Figure 10B ). In this example but not necessarily all examples, the obstructed spectrum is subtracted from the unobstructed spectrum. In this example, the difference 40 represents the spectrum of the audio that has been obstructed from the first region 16. The spectrum of the audio that has been obstructed from the first region 16 has a range 42.
[0222] The spectrum of the audio that has been obstructed from the first region 16 is transformed to the Figure 11A and 11B frequency filter 50. The filter applies a gain to the spectrum range 42 and the harmonics H outside of range 42.
[0223] In Figure 11A , the gain is a positive gain (>1) within range 42 and at harmonics H having a frequency lower than range 42, and otherwise a negative gain. It is a negative gain (<1) outside of range 42, including at harmonics H having a frequency higher than range 42. Thus, Figure 11A shows an amplification filter 50 that is configured to preferentially amplify the spectrum of the audio that has been obstructed from the first region 16 (and its low-frequency harmonics).
[0224] In Figure 11B , the gain is a negative gain (<1) within range 42 and at harmonics H having a frequency lower than range 42, and otherwise a positive gain (>1). It is a positive gain outside of range 42, including at harmonics H having a frequency higher than range 42. Thus,Figure 11B An amplified filter 50 is shown, which is configured to preferentially attenuate the spectrum (and its low-frequency harmonics) of the obstructed audio from the first region 16.
[0225] Therefore, it should be understood that the frequency-dependent filter 50 is configured as a space-dependent filter 50, which differentially amplifies or attenuates the audio sources 10 (and their low-frequency harmonics) at different spatial positions.
[0226] The frequency-dependent filter 50 is based on the spectral difference 40 between the obstructed audio input 22' and the unobstructed audio input 22, which is caused by the acoustic shadow of the obstacle 30. The obstacle 30 can be, for example, the hand 132 of the user 130, and the hand 132 provides at least a partial obstacle 30 between the obstructed microphone 20 and the first region 16.
[0227] The gain can be controlled by the user 130. For example, in some but not necessarily all examples, the user can select whether the filter 50 is an amplified filter or an attenuated filter. For example, in some but not necessarily all examples, the user 130 can determine the magnitude of the gain, e.g., the magnitude of amplification / attenuation. In some examples, the user input can be a gesture 134 of the hand 132 (see Figure 7 ). The gesture 134 can be part of or separate from the gesture for placing the hand 132 as the obstacle 30.
[0228] In this example, the frequency-dependent filter 50 can be, for example, based on the spectral difference of the audio input 22 from at least one microphone 20, which is caused by the acoustic shadow of the user's hand. The difference can be the change over time (time division) of one or more microphones 20. The difference can be the difference between microphones 20 at the same time (space division).
[0229] If the target audio at the first region 16 contains very low frequencies, then regarding Figures 10A to 10C and Figure 11A and Figure 11B The examples described may be useful. Very low frequencies may not be affected by the acoustic shadow caused by the hand.
[0230] For example, if the unobstructed spectrum ( Figure 10A ) has at the obstructed spectrum ( Figure 10B) For harmonic structures not present in , filter 50 can be extended to lower-frequency harmonics. For example, assume that harmonic frequencies 200 Hz, 300 Hz, 400 Hz, 500 Hz, 600 Hz, 700 Hz, 800 Hz, 900 Hz, and 1000 Hz are significantly attenuated by hand shadow. In this case, device 100 can determine that the target audio is likely a harmonic sound with a fundamental frequency of 100 Hz. Even though the frequency 100 Hz is not affected by hand shadow, the system can still select this frequency to be included in frequency-dependent filter 50.
[0231] The term "frequency" can refer to a single frequency or a frequency band with a certain width.
[0232] References to "computer-readable storage media", "computer program products", "tangibly embodied computer programs", etc., or "controllers", "computers", "processors", etc. should be understood to cover not only computers with different architectures (such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures), but also dedicated circuits, such as field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), signal processing devices, and other processing circuitry. References to computer programs, instructions, code, etc. should be understood to cover software for programmable processors or firmware, such as programmable content of hardware devices, whether instructions for a processor or configuration settings of a fixed-function device, gate arrays, or programmable logic devices, etc.
[0233] As used in this application, the term "circuitry" can refer to one or more or all of the following:
[0234] (a) Pure hardware circuitry implementation (such as implementation using only analog and / or digital circuitry), and
[0235] (b) A combination of hardware circuitry and software, such as (where applicable):
[0236] (i) A combination of (multiple) analog and / or digital hardware circuitry and software / firmware, and
[0237] (ii) Any part of (multiple) hardware processors (including (multiple) digital signal processors), software, and memory with software that work together to cause a device (such as a mobile phone or server) to perform various functions, and
[0238] (c) (Multiple) hardware circuitry and / or (multiple) processors, such as (multiple) microprocessors or a part of (multiple) microprocessors, which require software (e.g., firmware) to operate, but the software may not be present when not in operation.
[0239] The definition of circuitry applies to all uses of the term in this application, including in any claims. As another example, as used in this application, the term circuitry also encompasses implementations of only hardware circuits or processors and their (or their) accompanying software and / or firmware. For example, if applicable to a particular claim element, the term circuitry also encompasses a baseband integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or network device.
[0240] The boxes shown in the figures can represent steps in a method and / or segments of code in a computer program 76. The recitation of a particular order of the boxes does not necessarily indicate that the boxes have a required or preferred order, and the order and arrangement of the boxes can be changed. Additionally, some blocks may be omitted.
[0241] Where a structural feature has been described, the structural feature can be replaced with a component that performs one or more functions of the structural feature, whether the function or functions are explicitly or implicitly described.
[0242] The recording of data can include only temporary records, or it can include permanent records, or it can include both temporary and permanent records. Temporary records represent data that is recorded temporarily. This can occur, for example, during sensing or image capture, at dynamic memory, or at buffers such as circular buffers, registers, caches, etc. Permanent records represent data in the form of an addressable data structure that can be retrieved from an addressable memory space and thus can be stored and retrieved until deleted or overwritten, although long-term storage may or may not occur. The use of the term "capture" in relation to an image or audio relates to the temporary recording of data. The use of the term "record" or "store" in relation to an image or audio relates to the permanent recording of data.
[0243] The above examples can be used as enabling components for the following:
[0244] Automotive systems; telecommunications systems; electronic systems, including consumer electronics; distributed computing systems; media systems for generating or rendering media content, including audio, visual, and audiovisual content, as well as mixed, mediated, virtual, and / or augmented reality; personal systems, including personal health systems or personal fitness systems; navigation systems; user interfaces, also known as human-machine interfaces; networks, including cellular, non-cellular, and optical networks; ad hoc networks; the Internet; the Internet of Things; virtualized networks; and related software and services.
[0245] The term "comprising" as used in this document has an inclusive rather than an exclusive meaning. That is, any reference to X that comprises Y means that X can include only one Y or can include more than one Y. If it is intended to use "comprising" in an exclusive sense, it will be explicitly stated in the context by referring to "comprising only one..." or by using "consisting of".
[0246] In this specification, various examples are referred to. The description of a feature or function in relation to an example indicates that the feature or function exists in that example. The use of the terms "example" or "for example" or "may" or "can" in the text indicates that, whether or not explicitly stated, the feature or function exists at least in the described example, whether or not described as an example, and they can but do not necessarily exist in some or all other examples. Thus, "example", "for example", "may" or "can" refer to a particular instance within a class of examples. The properties of that instance can be properties of only that instance, or properties of the class, or properties of a subclass of the class that includes some but not all instances of the class. Thus, features described with reference to one example rather than another are implicitly disclosed as being usable, where possible, as part of a working combination in that other example, but do not necessarily have to be used in that other example.
[0247] Although examples have been described in the preceding paragraphs with reference to various examples, it should be understood that the examples given can be modified without departing from the scope of the claims.
[0248] The features described in the foregoing description can be used in other combinations than those explicitly described above.
[0249] Although functions have been described with reference to certain features, these functions can be performed by other features, whether or not described.
[0250] Although features have been described with reference to certain examples, these features can also exist in other examples, whether or not described.
[0251] The term "a" or "the" as used in this document has an inclusive rather than an exclusive meaning. That is, any reference to X that includes a / the Y means that X can include only one Y or can include more than one Y, unless the context clearly indicates the contrary. If it is intended to use "a" or "the" in an exclusive sense, it will be explicitly stated in the context. In some cases, "at least one" or "one or more" can be used to emphasize the inclusive meaning, but the absence of these terms should not be taken as inferring any exclusive meaning.
[0252] The presence of a feature (or combination of features) in a claim is a reference to that feature or (combination of features) itself, as well as a reference to features (equivalent features) that achieve substantially the same technical effect. Equivalent features include, for example, features that are variants and achieve substantially the same result in substantially the same way. Equivalent features include, for example, features that perform substantially the same function in substantially the same way to achieve substantially the same result.
[0253] In this specification, various examples are referred to, and adjectives or adjective phrases are used to describe the features of the examples. Such a description of a characteristic associated with an example indicates that the characteristic exists exactly as described in some examples and exists substantially as described in other examples.
[0254] Although the foregoing specification has sought to draw attention to those features that are considered important, it should be understood that the applicant may seek protection by means of claims for any patentable feature or combination of features mentioned above and / or shown in the drawings, whether or not emphasized.
Claims
1. An apparatus for audio processing, comprising: Components for: Detecting a user input indicating the presence of at least a partially user-controlled obstacle; Receiving a blocked audio input from the at least one blocked microphone when the user provides the at least partially user-controlled obstacle between the at least one blocked microphone and a first region; Receiving an unblocked audio input from the at least one unblocked microphone when the user does not provide the at least partially user-controlled obstacle between the at least one unblocked microphone and the first region; Comparing the blocked audio input and the unblocked audio input; Creating a frequency-dependent filter based on the comparison; Filtering an audio input received from the at least one microphone after removing the at least partially user-controlled obstacle to create a filtered audio that amplifies or attenuates an audio source in the first region; And A camera, and Wherein the component for detecting a user input indicating the presence of at least a partially user-controlled obstacle includes a component for processing an output from the camera to identify the presence of a hand and movement of the hand as a user gesture.
2. The apparatus according to claim 1, wherein the frequency-dependent filter is configured as a spatially dependent filter that differentially amplifies or attenuates audio sources at different spatial positions.
3. The apparatus according to claim 1 or 2, configured to detect a spatially specific gesture as the user input, the spatially specific gesture providing a spatially related at least partially user-controlled obstacle to audio reaching the at least one blocked microphone.
4. The apparatus according to claim 1 or 2, wherein the frequency-dependent filter is based on a spectral difference between the audio inputs between the at least one blocked microphone and the at least one unblocked microphone, the spectral difference being caused by an acoustic shadow of the at least partially user-controlled obstacle between the at least one blocked microphone and the first region.
5. The apparatus according to claim 1 or 2, comprising components for: Spectral analysis of the blocked audio input; Spectral analysis of the unblocked audio input; Generating the frequency-dependent filter based on a difference between the spectral analyses.
6. The apparatus according to claim 1 or 2, wherein the frequency-dependent filter selectively provides gain to spectral components corresponding to spectral components of the blocked audio signal.
7. The apparatus according to claim 1 or 2, wherein the frequency-dependent filter selectively provides gain to harmonics of spectral components having a harmonic structure, the harmonics of the spectral components being attenuated by the at least partially user-controlled obstacle, wherein the harmonics of the spectral components have harmonic frequencies lower than a predetermined harmonic frequency.
8. The apparatus according to claim 6, wherein the gain is controlled by the user.
9. The apparatus according to claim 1 or 2, wherein the frequency-dependent filter is a time-varying filter configured to decay over time.
10. The apparatus according to claim 1 or 2, configured to prompt the user to repeat creation of the frequency-dependent filter, including: When a user provides at least a partial obstacle between at least one blocked microphone and a first region, receive blocked audio input from the at least one blocked microphone; When the user does not provide the at least partial obstacle between at least one unblocked microphone and the first region, receive unblocked audio input from the at least one unblocked microphone; Compare the blocked audio input and the unblocked audio input; Create a frequency-dependent filter based on the comparison; After removing the at least partial user-controlled obstacle, filter the audio input received from at least one microphone to create filtered audio that amplifies or attenuates an audio source in the first region.
11. A method of audio processing, comprising: When a user provides at least a partial user-controlled obstacle between at least one microphone and a first region, receive blocked audio input from at least one blocked microphone; When the user does not provide the at least partial user-controlled obstacle between the at least one microphone and the first region, receive unblocked audio input from at least one unblocked microphone; Compare the blocked audio input and the unblocked audio input; Create a frequency-dependent filter based on the comparison; After removing the at least partial user-controlled obstacle, filter the audio input received from the at least one microphone to create filtered audio that amplifies or attenuates an audio source in the first region; and Detect a user input indicating the presence of at least a partial user-controlled obstacle, wherein detecting the user input includes processing an output from a camera to identify the presence of a hand and movement of the hand as a user gesture.
12. The method according to claim 11, wherein the frequency-dependent filter is based on a spectral difference between the audio inputs from the blocked microphone and the unblocked microphone, the spectral difference being caused by an acoustic shadow of the user's hand.
13. The method according to claim 11 or 12, wherein the frequency-dependent filter provides a frequency-dependent gain, and the gain is controlled by the user to attenuate or amplify.
14. A computer program product that, when run on at least one processor, the computer program performs: Detect a user input indicating the presence of at least a partial user-controlled obstacle; When a user provides the at least partial user-controlled obstacle between at least one blocked microphone and a first region, receive blocked audio input from the at least one blocked microphone; When the user does not provide the at least partial user-controlled obstacle between at least one unblocked microphone and the first region, receive unblocked audio input from the at least one unblocked microphone; Compare the blocked audio input and the unblocked audio input; Create a frequency-dependent filter based on the comparison; After removing the at least partial user-controlled obstacle, filter the audio input received from at least one microphone to create filtered audio that amplifies or attenuates an audio source in the first region, and Wherein detecting user input indicating the presence of an obstacle at least partially under user control includes processing the output from a camera to identify the presence of a hand and movement of the hand as a user gesture.
Citation Information
Patent Citations
Apparatus, method and computer program for determining microphone occlusion
CN117616781A
Hearing aid device and hearing aid method
US20120057733A1