Speaker audio masking
Patent Information
- Authority / Receiving Office
- ES · ES
- Patent Type
- Patents
- Current Assignee / Owner
- AUDIO MOBIL ELEKTRONIK GMBH
- Filing Date
- 2022-10-18
- Publication Date
- 2026-07-14
AI Technical Summary
Unwanted eavesdropping in zone-based audio systems, such as in vehicles or public transportation, is a challenge as existing solutions like playing loud noise increase noise levels and disturb others, making them undesirable.
A method and device that generate a masking signal by swapping spectral bands of a captured speech signal and adding a broadband noise signal, optionally with distraction signals, to reduce speech intelligibility without significantly increasing overall sound level.
Effectively reduces speech intelligibility in unintended listening zones while maintaining low overall sound pressure and comfort, providing private acoustic zones for speakers.
Smart Images

Figure 00000014_0000 
Figure 00000014_0001 
Figure 00000015_0000
Abstract
Description
[0001] The present disclosure relates to the generation of a masking signal for speech in a zone-based audio system.
[0002] Modern communication technologies and their ever-increasing coverage enable communication to take place almost anywhere, for example, in the form of telephone conversations. In public spaces, other people can often overhear such conversations and understand their content. This is particularly problematic when the conversations are confidential, private, or business-related. Such a scenario exists on public transportation, such as trains or airplanes, but also in private vehicles, such as taxis or rented limousines. In these cases, in addition to the speaker, other people are in fixed positions, for example, in assigned seats. Often, these seats have an associated audio system or at least components thereof.For example, speakers for individual playback of audio content may be provided in these seats, for example integrated into neck rests, which is also referred to as a zone-based audio system.
[0003] Besides telephone conversations, the problem of unwanted eavesdropping can also occur in conversations between people. For example, two passengers in the back of a taxi might be discussing a confidential topic, and the driver might not want to overhear them.
[0004] It is known from current technology that unwanted eavesdropping can be reduced by playing loud noise. However, this increases the noise level for everyone involved and is perceived as an unpleasant disturbance that can also affect attention and reaction time, which is particularly undesirable in road traffic.
[0005] This document addresses the technical challenge of generating a masking signal in a zone-based audio system that reduces unwanted overhearing of a conversation without causing any unpleasant disturbance.
[0006] The problem is solved by the features of the independent claims. Advantageous embodiments are described in the dependent claims. <1a>
[0007] According to a first aspect, a method for masking a speech signal in a zone-based audio system is disclosed. The method comprises capturing a speech signal to be masked in an audio zone, for example, by means of one or more conveniently placed microphones, which may be located, for instance, in the headrest of a seat. The speech signal may originate from the local speaker of a telephone conversation or belong to a conversation between people present. The captured speech signal is then transformed into spectral bands, which can be done, for example, using an FFT and Mel filters. Furthermore, the method involves swapping spectral values of at least two spectral bands, thereby altering the spectral structure of the speech signal without changing its overall energy content. Subsequently, a (preferably broadband) noise signal is generated based on the swapped spectral values.The generated noise signal indicates
[0008] US Patent 2012 / 016665 A1 discloses a device for generating a masking signal, wherein a CPU analyzes the speech utterance rate of a received audio signal. The CPU then copies the received audio signal into a multitude of audio signals and performs the following processing for each of the audio signals. Specifically, the CPU divides each of the audio signals into frames based on a frame length determined by the speech utterance rate. A reversal process is performed on each of the frames to replace one waveform of the frame with an inverted waveform, and a windowing process is performed to achieve smooth transitions between the frames. Subsequently, the CPU randomly reorders the sequence of the frames and mixes the multiple audio signals to generate a masking audio signal.
[0009] German patent application DE 10 2014 214052 A1 discloses a method and a corresponding device for generating a shielded listening zone within a vehicle. It describes a method for masking a target speech signal. The method comprises identifying the target speech signal, which is reproduced in a first listening zone of the vehicle, and generating a first masking sound signal based on the identified target speech signal. The method also includes binauralizing the first masking sound signal to generate a binaural masking sound signal. Furthermore, the method includes reproducing the binaural masking sound signal via at least two loudspeakers in a second listening zone of the vehicle, where the target speech signal is to be masked.
[0010] In "Aircraft noise and speech intelligibility in an outdoor living space" a "Partial Loudness Model" is developed that predicts the loudness of a signal in the presence of a masking sound signal, taking into account masking across spectral bands and the effects of temporal masking of time-varying noise.
[0011] While it exhibits a certain similarity to the spectrum of the speech signal, it does not perfectly match it, as the spectral structure of the speech signal is no longer fully preserved due to the band swapping. Such a noise signal with a similar but not identical spectrum to the speech signal is well-suited as a masking signal for the speech signal. It should also be noted that any number of bands can be swapped (e.g., all of them), with more band swaps resulting in greater variation in the noise spectrum. Finally, the noise signal is output as a masking signal in a different audio zone with minimal energy input to make it more difficult for a person located at that listening position to overhear the conversation by reducing speech intelligibility for them.
[0012] Generating a noise signal based on swapped spectral values can involve creating a broadband noise signal, for example, using a noise generator, and transforming the generated noise signal into the frequency domain. Furthermore, the frequency representation of the noise signal can be multiplied by a frequency representation of the speech signal, taking the swapped spectral values into account. This frequency-domain multiplication produces a noise spectrum that essentially corresponds to that of the speech signal after the spectral bands have been swapped—that is, it is similar to, but not identical to, the speech spectrum. A similar effect can also be achieved through convolution in the time domain.
[0013] The frequency representation of the speech signal can be generated by interpolating the spectral values of the bands (for example, in the mel range) after swapping the spectral values. This interpolation generates the necessary values at the frequency reference points for multiplication with the noise spectrum from the (relatively few) spectral values of the bands.
[0014] The method can further involve estimating a background noise spectrum (preferably at the listening position) and comparing spectral values of the speech signal with the background noise spectrum. The comparison of spectral values preferably (but not necessarily) takes place within the spectral bands (e.g., mel bands), which means that the background noise spectrum must also be represented in these spectral bands. Furthermore, only spectral values of the speech signal that are larger than the corresponding spectral values of the background noise spectrum (or bear a predetermined ratio to them) can be considered for further processing (e.g., the interpolation mentioned above). Spectral components of the speech signal that are already masked by the background noise do not need to be considered for generating the masking signal and can be suppressed (e.g., by setting them to zero).Background noise can be accounted for both before and after the swapping of spectral values. In the former case, the spectral bands being compared still match exactly, and the background noise is correctly accounted for. In the latter case, swapping bands and attenuating low-energy bands in the speech signal introduces an additional variation into the noise spectrum, which can lead to increased masking. This allows for a masking signal adapted to the background or environment, which can be output to the listener's audio field with minimal energy input.
[0015] The transformation of the captured speech signal into spectral bands can be performed for blocks of the speech signal and using a Mel filter bank. Optionally, it is possible to perform time smoothing of the spectral values for the Mel bands, e.g., in the form of a moving average.
[0016] In a further embodiment of the invention, the noise signal can be spatially represented during output by means of a multi-channel (i.e., at least two-channel) playback. For this purpose, a multi-channel representation of the masking signal can be generated, enabling spatial reproduction of the masking signal. For two-channel systems, this can preferably be achieved by multiplication with binaural spectra of an acoustic transfer function. The spatial reproduction enhances the effect of the masking signal on concealing speech at the listening position, particularly when the noise signal is output spatially in the other audio zone in such a way that it appears to originate from the direction of the speaker of the speech signal to be masked.
[0017] In addition to the masking signal described above, which is based on a broadband noise signal adapted to the speech signal, a further component can be generated for the masking signal and output together to the listener in the second audio zone. For this purpose, the method can involve determining a point in the speech signal relevant to speech intelligibility (e.g., the presence of consonants) and generating a suitable distraction signal for that specific point in time. The distraction signal can then be output at that specific point in time as a further masking signal in the other audio zone, resulting in a localized additional obfuscation (masking) of the speech content during speech onsets. Since the distraction signal is only output at specific relevant points in time, it does not significantly increase the overall sound level and does not cause any significant impairment.
[0018] The relevant time for speech intelligibility can be determined based on extreme values (e.g., local maxima, onsets) of a spectral function of the speech signal. This spectral function is determined by summing spectral values along the frequency axis. The spectral values can be smoothed beforehand in the temporal and / or frequency direction. After summing the spectral values along the frequency axis, the sums can optionally be logarithmized. To generate local maxima for detecting relevant time points, the (optionally logarithmized) sums can be time-differentiated.
[0019] Furthermore, the time points relevant for speech intelligibility can be verified using parameters of the speech signal, such as zero-crossing rate, short-term energy, and / or spectral center. It is also possible to consider temporal constraints for extreme values, such as requiring them to have a predetermined minimum time interval.
[0020] The distraction signal for a specific time can then be randomly selected from a set of predefined distraction signals. These can be stored in a memory for selection. It has proven advantageous to match the distraction signal to the speech signal with respect to its spectral characteristics and / or energy. In this way, the spectral center of the distraction signal can be aligned with the spectral center of the corresponding speech segment at the specific time, for example, by means of single-sideband modulation. A speech segment with a high spectral center can thus be masked with a distraction signal that also has a high spectral center (possibly even the same spectral center), leading to greater masking effectiveness.The energy of the distraction signal can also be adjusted to match the energy of the speech segment in order to avoid creating a masking signal that is too loud and excessively disruptive.
[0021] In a further embodiment of the invention, the distraction signal can be represented during output by means of a multi-channel spatial reproduction, preferably by multiplication with binaural spectra of an acoustic transfer function, thereby generating a multi-channel (at least two-channel) representation of the distraction signal, which enables a spatial reproduction of the distraction signal. This spatial reproduction enhances the effect of the distraction signal on masking speech at the listening position, particularly when the distraction signal is output spatially in the other audio zone in such a way that it appears to originate from a random direction and / or near the listener's head in that other audio zone. This spatialization reduces the distinguishability of the speech and distraction signals, or makes it more difficult to overhear the speech signal due to the distraction signal, and thus reduces the energy required for the distraction signal.
[0022] The processing of the speech signal and the generation of a masking signal described above are preferably carried out in the digital domain. This requires steps not described in detail, such as an analog-to-digital conversion and a digital-to-analog conversion, which, however, will be self-evident to those skilled in the art after studying the present disclosure. Furthermore, the above method can be implemented wholly or partially by means of a programmable device, which in particular includes a digital signal processor and the necessary analog-to-digital converters.
[0023] According to a further aspect of the invention, a device for generating a masking signal in a zone-based audio system is proposed, which receives a speech signal to be masked and generates the masking signal based on the speech signal. The device comprises means for transforming the acquired speech signal into spectral bands; means for interchanging spectral values of at least two spectral bands; and means for generating a noise signal as a masking signal based on the interchanged spectral values.
[0024] The embodiments of the method described above can also be applied to this device. The device can further comprise: means for determining a point in time relevant to speech intelligibility within the speech signal; means for generating a distraction signal for the relevant point in time; and means for adding the noise signal and the distraction signal and outputting the sum signal as a masking signal.
[0025] In a further embodiment of the device, it also includes means for generating a multi-channel representation of the masking signal, which enables a spatial reproduction of the masking signal.
[0026] According to a further aspect of the invention, a zone-based audio system with a plurality of audio zones is disclosed, wherein at least one audio zone has a microphone for capturing a speech signal and another audio zone has at least one loudspeaker. The microphone and loudspeaker can be arranged in the headrests of seats for vehicle occupants. It is also possible for both audio zones to have a microphone and loudspeaker. The audio system includes a device, as shown above, for generating a masking signal, which receives a speech signal from a microphone of one audio zone and sends the masking signal to the loudspeaker(s) of the other audio zone.
[0027] Another aspect of the present disclosure relates to the generation of a distraction signal as a masking signal, independent of the aforementioned noise signal, as described above. A corresponding method for masking a speech signal in a zone-based audio system comprises: capturing a speech signal to be masked in one audio zone; determining a point in time relevant to speech intelligibility within the speech signal; generating a distraction signal for the determined point in time, wherein the distraction signal may be adapted to the speech signal with respect to a spectral characteristic and / or its energy; and outputting the distraction signal at the determined point in time as a masking signal in the other audio zone. The possible embodiments of the method correspond to the embodiments described above in combination with the generated noise signal.
[0028] A corresponding device for generating a distraction signal as a masking signal in a zone-based audio system is also disclosed. This device receives a speech signal to be masked and generates the masking signal based on the speech signal. It includes means for determining a point in time relevant to speech intelligibility within the speech signal; means for generating a distraction signal for the relevant point in time, wherein the distraction signal may be adapted to the speech signal with respect to its spectral characteristics and / or energy; and means for outputting the distraction signal as a masking signal. Optionally, means for generating a multi-channel representation of the masking signal, enabling spatial reproduction of the masking signal, may be provided.
[0029] The features described above can be combined in many ways, even if such a combination is not explicitly mentioned. In particular, features described for a process can also be used for a corresponding device, and vice versa.
[0030] Exemplary embodiments of the invention will now be described in more detail with reference to the schematic drawing. These show: Fig. 1 schematically shows an example of a zone-based audio system; Fig. 2 schematically shows another example of a zone-based audio system; Fig. 3 schematically shows another example of a two-zone zone-based audio system; Fig. 4 schematically shows another example of a multi-zone zone-based audio system; Fig. 5 shows an example of a block diagram for generating a broadband masking signal to obscure speech; and Fig. 6 shows an example of a block diagram for generating a deflection signal to obscure speech.
[0031] The exemplary embodiments described below are not limiting and are purely illustrative. For the purpose of illustration, they include additional elements that are not essential to the invention. The scope of protection is to be determined solely by the accompanying claims.
[0032] The following examples allow vehicle occupants in any seating position to conduct undisturbed private conversations, such as phone calls with other people outside the vehicle. For this purpose, an audio masking signal is generated and transmitted to other vehicle occupants, disrupting their perception of the conversation and making it difficult, or ideally impossible, for them to overhear the private conversation. This creates a private space for the speaker, allowing them to conduct private conversations undisturbed, without the risk of other vehicle occupants overhearing confidential information. The conversation could be, for example, a phone call or a conversation between vehicle occupants.In the latter case, there are two speakers who alternately emit speech signals that other inmates should not understand if possible, while of course ensuring that the speech intelligibility between the two participants in the conversation is not impaired.
[0033] Similar scenarios arise more generally when individuals are located in acoustic zones or environments within a space, each served by separate acoustic playback devices. Such acoustic zones can exist, for example, in means of transport such as vehicles, trains, buses, airplanes, ferries, etc., where passengers are seated in seats equipped with individual acoustic playback devices. However, the proposed approach to creating private acoustic zones is not limited to these examples. It can be applied more generally to situations where individuals are located in specific positions within a space (e.g., in theater or cinema seats) and can be served by individual acoustic playback devices, and where there is a possibility of capturing the speech signals of a speaker whose speech should not be understood by the other occupants.
[0034] In one embodiment, a zone-based audio system is provided to create private acoustic zones at each passenger seat in a vehicle, or more generally, an acoustic environment. The individual components of the audio system are networked and can interact and exchange information / signals. Figure 1 Figure 1 schematically shows an example of such a zone-based audio system. A user or passenger is located in a seat 2 with a headrest 3, which has two loudspeakers 4 and two microphones 5.
[0035] Such a zone-based audio system has one, preferably at least two, loudspeakers 4 for the active acoustic reproduction of personal and individual audio signals, which should not be perceived, or only minimally perceived, by neighboring zones. The loudspeaker(s) 4 can be installed in the headrest 3, the seat 2 itself, or in the vehicle's headliner. The loudspeakers have a sufficiently optimized acoustic design and can be controlled via appropriate signal processing to minimize the acoustic impact on neighboring zones.
[0036] Furthermore, such an audio zone has the capability to record the speech of the occupant of the primary acoustic zone independently of the neighboring zones and the signals actively reproduced therein. For this purpose, one or more microphones 5 can be integrated into the seat 2 or the headrest 3, or placed in the immediate acoustic environment of the zone and the occupant, as shown in Figure 2 is shown schematically. Preferably, the microphones 5 are arranged to enable the best possible capture of the speech of the occupant making the phone call. If a microphone can be placed in the immediate vicinity of the speaker's mouth (like the middle microphone in Figure 2Generally, a single microphone is sufficient to capture the speaker's audio signals with adequate quality. For example, the microphone of a telephone headset can be used to record speech signals. Otherwise, two or more microphones are advantageous for capturing speech in order to record it better and, above all, more precisely using digital signal processing, as explained below.
[0037] The speaker's audio zone can have appropriate signal processing to record the primary occupant's speech signals as undisturbed as possible and unaffected by neighboring zones and prevailing disturbances in the environment (wind, rolling noise, ventilation, etc.).
[0038] The voice signal of the vehicle occupant making a phone call is thus captured at the seating position (either directly by a suitably positioned microphone or indirectly by means of one or more remote microphones with appropriate signal processing) and separated from any interfering signals, such as background noise.
[0039] From this speech signal, a masking signal, also referred to as a speech obfuscation signal, can be generated for an overheard passenger. In some embodiments, a broadband masking signal adapted to the speech to be obfuscated is generated for this passenger. Additionally or alternatively, distraction signals can also be generated at the individual speech onsets within the primary speaker's speech. These are short interference signals that are emitted at specific speech segments important for speech intelligibility and can also be adapted to the speech to be obfuscated. These distraction signals are emitted with temporal overlap with the speech segments relevant for speech intelligibility in order to reduce the information content for the listener and thus improve speech intelligibility.to impair their interpretation (informational masking) without significantly increasing the overall sound level.
[0040] Adapted to the specific local acoustic requirements, these masking signals can be played back spatially (multi-channel), creating a spatial perception of the masking signals. In this way, eavesdropping from the seating positions of the listening individuals can be avoided as effectively as possible.
[0041] The proposed approach ensures that the overall sound pressure level at the seating positions of the listening passengers increases only minimally and that the annoyance or impairment (annoyance) of the passengers is not increased, or that local listening comfort is maintained as much as possible, in contrast to an approach in which a loud background noise is simply emitted to mask the speech (energetic masking).
[0042] Figure 3Figure 1 illustrates the functionality and basic system structure of an embodiment for two audio zones. The speech signals of the occupant in the primary acoustic zone I are captured by microphones 5 located in the speaker's headrest 3 of that zone and subjected to a first digital signal processing A to record the speech signals of the primary occupant as undisturbed as possible and unaffected by neighboring zones and ambient disturbances (wind, road noise, ventilation, etc.). Alternatively, the microphone(s) 5 can also be located in front of the speaker, as shown in Figure 1. Figure 2This can be depicted, for example, in the rear part of the front passenger's headrest, or in the headliner, steering wheel, or dashboard. In the example shown, the eavesdropper is in the seat directly in front of the speaker, but this is not mandatory, and the eavesdropper can be located anywhere else within the vehicle.
[0043] The processed speech signals are then fed to a second signal processing unit B, which generates appropriate speech masking signals to reduce the speech intelligibility of the eavesdropping occupant. These speech masking signals are then output via loudspeakers 4' in the second acoustic zone II. These loudspeakers are located, for example, in the headrest 3' of the eavesdropping occupant to ensure the most direct and undisturbed reproduction of the speech masking signals possible. As mentioned previously, a speech masking signal can consist of a broadband masking signal adapted to the speech signal of the primary occupant and / or a distraction signal that begins at specific speech inflection points. In this way, acoustic zones can be designed to be so private that unwanted eavesdropping across the boundary of an acoustic zone is significantly hindered.
[0044] An alternative solution—similar to active noise cancellation—reduces the estimated speech signals at the respective listening and microphone locations by actively applying adaptive cancellation signals. However, since the listening position is easily variable in practice, and the listening and microphone locations are often several centimeters apart, this approach can only actively reduce speech signal components up to approximately 1.5 kHz. Because speech intelligibility is primarily dominated by consonants and thus signal components with frequencies above 2 kHz, this approach alone is insufficient and, at best, problematic. If the system is not properly calibrated (e.g., incorrectly adjusted to the head position), the cancellation signals can carry precisely the relevant private information and even amplify it, thus increasing rather than decreasing speech intelligibility.In contrast, the proposed approach is less sensitive to the exact head positions of the speaker and the listener and allows for a reduction in speech intelligibility even of higher frequency speech components such as consonants.
[0045] Due to the modularity of the proposed approach, implementation examples with multiple audio zones are also conceivable, such as in mass transport (railway, airplane, train) or other fields of application (entertainment, cinema, etc.). Figure 4 This schematically illustrates such a multi-zone approach using a multi-row vehicle as an example, in which six acoustic zones are provided. As before, loudspeakers and microphones are integrated into the passengers' headrests, although the microphones can also be positioned in other locations in front of the respective speakers to optimize speech signal capture. Similar to in Figure 3In this example, it is assumed that the speaker is sitting behind the unwanted eavesdropper (here, the driver). However, the speaking occupant's voice signals can be used in the same way to generate masking or concealment signals for occupants other than the driver, and even for multiple unwanted eavesdroppers. Naturally, the speaker could also be located in a different place within the vehicle than the one described in the example. Figure 4 The example shown illustrates this approach, which can be applied generally to all scenarios where a speaker's speech is captured and the generated speech obfuscation signals can be specifically output to the unwanted eavesdropper(s).
[0046] As mentioned at the beginning, the speech signals could be a telephone conversation that the speaker is having with an external person outside the room containing the acoustic zones. Alternatively, the conversation could also be taking place between people in the room, for example, between the person in Figure 4 The system detects the speaker and the occupant to their right. In this case, the zone-based audio system must perform the same signal processing for the second speaker as for the first speaker, so that their speech is also captured and processed to generate appropriate masking signals for the listener(s). When the two speakers alternate, only the current speaker needs to be identified and the masking signals associated with that speaker output. If both speakers speak simultaneously, both masking signals can be output at the same time.
[0047] The following describes the necessary signal processing steps for an exemplary use case. In this use case, a rear-left seated passenger is conducting a telephone conversation with a person outside the vehicle as the internal speaker. In addition to the internal speaker's voice, the voice of the external speaker, for example, output from the internal speaker's headrest speaker, can also be processed. For End Speaker Signal The speech to be masked is then recorded. This speech is then retouched or masked for the driver listening from the "front left" position. Of course, this is only one possible scenario, and the proposed methods can be applied generally to all possible configurations of speaker and listener positions.
[0048] The signal estimated using digital signal processing A say yesThe base parameter for the subsequent generation of the masking or obfuscation signal is provided for the speech signal to be masked. The speech signal to be masked can be the active internal speaker in the vehicle compartment and / or the external speaker outside. The obfuscation signal can be a broadband masking signal and / or distraction signals. These generated signals ( send to: out LS-Left & LS-RightThe signals are reproduced via the active neck support at the listening position. In some embodiments, both masking signals are generated, added, and reproduced together to amplify their effect on the listener and impair their intelligibility. The combination of the two masking signals creates a synergistic effect in reducing speech intelligibility. The continuous broadband masking signal generates background noise, and the volume (energy) of this signal can be reduced compared to outputting only a noise signal, resulting in a less disruptive effect. By selectively outputting the distracting signals at appropriate positions (speech onsets), the intelligibility of these speech segments (e.g.,(for consonants) are disrupted without significantly increasing the overall energy of the obfuscation signal or causing additional unpleasantness to the listener. It has even been found that the distraction signals are perceived as less unpleasant when presented together with the noise signal.
[0049] Figure 5 This shows a schematic block diagram for generating a broadband speech-signal-dependent mask. The input signal is the speech signal to be masked. say yes. The resulting two-channel output signals (out LS-Left & LS-Right) are sent to the active neck support at the listening position, possibly superimposed with distraction signals, and output to the listening person via loudspeakers attached to / in the neck support.
[0050] The following section describes in detail the signal processing steps for generating a broadband noise signal for speech masking, according to one exemplary implementation. It should be noted that not all steps are always necessary, and some steps can be performed in a different order, as experts in digital signal processing know. Furthermore, some calculations can be performed equivalently in either the frequency or time domain.
[0051] First, the speech signal say yes The signal is transformed into the frequency domain and smoothed both temporally and in the frequency direction. For this purpose, the speech signal is first processed in section 100. say yesThe signal is divided into blocks (for example, 512 samples at a sampling rate of fs = 44.1 kHz are arranged into blocks with a duration of 11.6 ms and 50% overlap). Subsequently, each signal block is transformed into the frequency domain in Section 105 using a Fourier transform with NFFT 1 = 1024 points.
[0052] In a further step (110), the Fourier spectra are filtered using a Mel filter bank with M = 24 bands – that is, the spectra are spectrally compressed by the Mel filter bank. The filter bank can consist of overlapping bands with a triangular frequency response. The center frequencies of the bands are equidistant across the Mel scale. The lowest frequency band of the filter bank starts at 0 Hz, and the highest frequency band ends at half the sampling rate (fs). For each band of the filter bank, a short-term energy value (RMS level or specific loudness response of the individual Mel bands) is calculated for each signal block in section 115 of the block diagram. These short-term energy values are time-averaged over MA = 120 blocks in section 120 using a moving average (120 blocks correspond to approximately 700 ms).
[0053] In exemplary implementations described in Section 125, these dynamic loudness profiles in the immediate frequency environment are interchanged (scrambled). For this purpose, the loudness values of the bands are swapped according to the following table, where the assignment of the band "in" is determined by the corresponding position in the row below, "out". For example, the loudness value of band number 2 is assigned to band number 4, and the value of band 4 is assigned to band 5, whose value is assigned to band 3, and so on. This results in swaps of loudness values with adjacent or next-but-one bands; that is, the difference between a Mel band and a swapped band is, in this example, a maximum of two Mel bands. Of course, the table shown is only one possible example of band swapping, and other implementations are possible.
[0054] The proposed band swapping technique "scrambles" the loudness values, creating a degree of "disorder" in the loudness distribution for a given speech segment. This alters the description of its spectral energy, or loudness distribution, without changing the overall energy or loudness of the speech segment. For example, a particularly high energy content in one band is shifted to another, or low energy (loudness) in one band is transformed into an adjacent band. It has been shown that this redistribution of energy to neighboring bands can generate a particularly effective broadband noise signal, which reduces the intelligibility of the corresponding speech segment more significantly than without band swapping.By swapping / rotating the sequence of the bins of the time-dynamic profiles of the masking bands, the transmission of speech information within the noise signal is avoided. If one were to capture the speech energy in frequency bands (e.g., mel bands as described above) and directly modulate these time-dynamic energy profiles onto a noise signal, also divided into equal frequency bands, the speech content would become audible—and even more so if narrow frequency bands are used. This effect is significantly reduced by swapping the loudness values within the bands.
[0055] The potentially reversed dynamic loudness curves can be adjusted based on the current background spectra (including all interfering noise) in section 130 of the block diagram to assess background noise and the environmental situation. For this purpose, the background noise is captured, for example, at the listening position, and the background spectra are determined by frequency transformation and temporal and frequency averaging, similar to the process used for the speech signal. Preferably, a microphone positioned at the listening position is used for this purpose. Alternatively, microphones located elsewhere (but as close as possible to the listening position) can also be used to capture the background noise at the listening position. Only those bands of the speech signal that lie above the background spectrum need to be considered when generating the masking signal.Speech bands whose energy is below that of the corresponding background noise band can be disregarded, as they are irrelevant for speech intelligibility or are already masked by the background noise. This can be achieved, for example, by setting the loudness value of such speech bands to zero. In other words, if a frequency band is already masked by strong background noise, no additional masking signal is generated in that frequency band. Thus, the system determines, based on the situation, which signal components of the broadband masking noise are used to obscure the speech.
[0056] In section 135, the resulting audibility thresholds (frequency axis sampled at 24 frequencies corresponding to the 24 center frequencies of the Mel filter bank) are interpolated at all frequency points of the Fourier transform. This interpolation generates a spectral value for the speech signal across the entire frequency range of the Fourier transform, for example, 1024 values for the aforementioned Fourier transform with NFFT 1 = 1024 points.
[0057] Finally, in section 155, the frequency reference points are multiplied (or convolutioned in the time domain) of the frequency values thus generated by a noise spectrum. This can be obtained by a noise generator (not shown), whose noise signal is analogous to the speech signal. say yesThe processing is carried out by block segmentation 145 and Fourier transformation 150 with identical dimensions. In this way, a broadband noise signal is generated as a masking signal with a similar frequency characteristic (apart from the interchange and zeroing of sections 125 and 130) to the speech signal. Alternatively, the masking signal can also be generated in the time domain by convolution of the noise signal with the spectral values of the speech signal processed as described above (see sections 100 to 135), which have been transformed back into the time domain. By switching between the frequency and time domains, different frequency resolutions or time durations can be used in the various processing steps. Alternatively, it is also possible to perform the entire processing in the frequency domain. For each block of the speech signal, a broadband noise spectrum adapted to the speech segment of the block is thus generated.
[0058] In exemplary embodiments, Section 160 describes a spatial processing procedure involving point-by-point multiplication of the frequency support points (or convolution in the time domain) with binaural spectra of an acoustic transfer function that corresponds to the source direction of the speaker (or the dominant direction of the energy center of the speech signal to be masked) from the perspective of the listening person. The source direction of the speaker is known from the spatial arrangement of the acoustic zones. In the example in Figure 4 In the example shown, the speaker's source direction is directly behind the listener. In embodiments with spatial orientation of the masking signal, multi-channel playback (e.g., using two loudspeakers) is required. Otherwise, single-channel playback is sufficient, preferably also using two loudspeakers positioned in the headrest of the listener.
[0059] The broadband masking signal can thus be spatially reproduced and adapted to the target direction of the direct signal or the prominently perceived direction of the speaker. Due to the binaural loudness addition, this results in significantly improved masking with lower level excesses of the masking noise.
[0060] In Section 165, an inverse transform (IFFT) of the two spectra resulting from spatial playback (per block) into the time domain is performed, and the blocks are superimposed using the overlap-add method (see Section 170). It is noted that this results in a multi-channel signal for spatial playback, which can be played back, for example, via stereo playback. If the previous steps have already been performed in the time domain, the inverse transform and the superimposition of the blocks are, of course, unnecessary.
[0061] The resulting time signals are sent to the respective active neck rest of the listener. In embodiments where distraction signals are also generated, the masking signals can be summed with the distraction signals before being output via the neck rest's loudspeakers.
[0062] As mentioned previously, signal processing can be performed partially in the frequency domain or in the time domain, and it is also possible to perform all processing in the frequency domain. The specific values mentioned above are only examples of a possible configuration and can be modified in many ways. For example, a frequency resolution of the FFT transformation with fewer than 1024 points or a division of the Mel filters with more or fewer than 24 filters is possible. It is also possible that the frequency transformation of the noise signal is performed with a different block size and / or FFT configuration than that of the speech signal. In this case, the interpolation in Section 135 would have to be adjusted accordingly to generate suitable frequency values.In a further variation, the block-wise calculated masking noises are first transformed back into the time domain after interpolation and then again into the frequency domain to take spatialization into account – possibly with a different spectral resolution. Those skilled in the art will recognize such variations of the inventive procedure for generating a broadband speech-signal-dependent masking signal after studying the present disclosure.
[0063] In exemplary implementations, short-duration deflection signals are used instead of masking noise. These signals are adapted, in terms of timing and / or frequency, to sections of the speech signal that are particularly relevant for intelligibility. An example of generating such deflection signals is described below. Figure 6The diagram schematically shows an example of a block diagram for generating speech-signal-dependent distraction signals. The eavesdropper is distracted at signal-dependent, defined times. For this purpose, the critical times (ti,distract) are determined based on three information parameters in the speech signal: spectral centroid "SC" (corresponding approximately to the pitch), short-term energy "RMS" (corresponding approximately to the loudness), and number of zero crossings "ZCR" (for distinguishing speech signal from background noise).
[0064] A digital memory contains a series of pre-selected distraction signals (e.g., bird calls, chirps, etc.) with corresponding parameters (SC and RMS), determined through additional pre-analysis. Suitable distraction signals preferably have the following characteristics: Firstly, they are natural signals familiar to the listener from other situations / daily life and therefore unrelated to the signal and context to be masked. Secondly, they are characterized by being acoustically distinctive signals of short duration and exhibiting a broadband spectrum. Other examples of such signals include the sound of water dripping or rippling, or brief gusts of wind. Typically, the distraction signals are longer than the relevant speech segments (e.g., consonants) and completely mask them.It is also possible to store distraction signals of different lengths and select them to match the duration of the current critical time.
[0065] A distraction signal is selected and adjusted in time and frequency to the current speech segment. This adjusted distraction signal can then be reproduced from a virtual spatial position for the listener. For spatialization (BRTF), short impulse responses (256 points) can be used to simulate the outer ear transfer function, ensuring that these distraction signals are localized as close and prominently as possible to the listener's head, thus achieving a strong distraction effect. Multichannel playback (e.g., in stereo) is required for spatial reproduction.
[0066] The following describes in detail the signal processing steps for generating discrete, spatially distributed, short deflection signals according to one embodiment. It should be noted that not all steps are always necessary, and some steps can be performed in a different order, as experts will recognize. Furthermore, some calculations can be performed equivalently in the frequency domain or the time domain. Some of the processing steps correspond to those for generating broadband masking signals and therefore do not need to be repeated in embodiments that use both signal types for speech masking.
[0067] Section 200 contains the speech signal say yes divided into blocks (BlockLength = 512 Samples, fs = 44.1kHz) with a duration of 11.6 ms and 50% overlap (HopSize = 256) (see section 100).
[0068] From these blocks XBuffer n (m), where n = block index and m = time sample, the number of zero crossings (zero-crossing rate, ZCR) per signal block is determined in Section 205. This can be done using the following formula: ZCR n = 0.5 ∗ ∑ m = 1 m BlockLength − 1 sgn XBuffer n m + 1 − sgn XBuffer n m
[0069] In section 210, each signal block is subjected to a Fourier transform with NFFT 2 = 1024 points (see section 105).
[0070] From these spectra S(k,n) with k = frequency index and n = block index, two further parameters are calculated in sections 215 and 220: the short-term energy (RMS) and the spectral centroid (SC): RMS n = ∑ k S k n 2 SC n = ∑ k = 1 NFFT 2 2 + 1 k ∗ S k n ∑ k = 1 NFFT 2 2 + 1 S k n
[0071] The short-term energy (RMS) and zero-crossing rate (ZCR) curves can be further filtered using signal-dependent thresholds, and areas that do not meet these thresholds can be hidden (e.g., set to zero). The thresholds can be chosen, for example, so that a certain percentage of the signal values lie above or below them.
[0072] Each spectrum is spectrally smoothed in both directions in Section 225 using a recursive time-discrete first-order filter: H(z) = Bs(z) / As(z), where Bs = 0.3 and As(z) = 1 - (Bs-1)*z -1< (= acausal, zero-phase second-order filter).
[0073] The resulting spectra are smoothed in time in section 230 using a recursive first-order discrete-time filter: H(z) = Bt(z) / At(z), where Bt = 0.3 and At(z) = 1- (Bt-1)*z -1<.
[0074] For the detection of speech intelligibility-relevant sections (onsets) of the speech signal (onset detection), an onset detection function is first determined in Section 235. For this purpose, the spectrally and time-averaged spectra are added along the frequency axis. The resulting signal is logarithmically and time-differentiated, with negative values being set to zero. Prior to the logarithm, regularization (e.g., adding a small number at each frequency point) can be performed to avoid zero values.
[0075] This onset detection function is examined for local maxima, which must be at least a predefined number of blocks apart. The maxima found in this way can be further filtered using a signal-dependent threshold, so that only particularly pronounced maxima remain. Local maxima of the onset detection function determined in this way are candidates for perceptually relevant sections of the speech signal that are to be selectively disrupted by a distraction signal.
[0076] In exemplary implementations, the maxima of the onset detection function determined in this way are checked for plausibility by a logic unit in Section 240 using the parameters ZCR, RMS, and SC. Only if these values lie within a defined range are these maxima considered relevant, critical time points. ten,distractThis can be achieved, for example, by requiring that the values of RMS, SC, and / or ZCR must meet certain logical conditions at the times of the determined maxima of the onset detection function (e.g., RMS > X1; X2). <SC<X3; ZCR> X4 with predefined thresholds X1 to X4). In exemplary implementations, for example, only those maxima are considered that lie in time intervals that satisfy the aforementioned filter conditions for RMS and ZCR (i.e., not in hidden regions). The condition that ZCR and RMS must simultaneously satisfy certain threshold conditions can also be used to filter the course of SC by preserving the values of SC when the threshold conditions are met and interpolating or extrapolating intermediate values, thus creating the function SC int.
[0077] At the determined times ten,distractFrom a selection of N deflection signals stored digitally in memory, one is randomly selected at a time (using section 245). Additional metadata for these deflection signals is stored in memory 250: SC and RMS values.
[0078] The selected deflection signal is divided into blocks in Section 255 (see above with BlockLength 2 and Hopsize = BlockLength 2 or Overlap = 0) and then Fourier-transformed using NFFT 2 points in Section 260. The parameters of this frequency transformation can be different and independent of the above procedure for the speech signal to be masked. Alternatively, the frequency representation of a deflection signal could also be stored directly in the frequency domain.
[0079] The resulting spectra can be described in Section 265 as signal-dependent. say yes at the respective time ten,distract based on the SCParameter ratios in the frequency range (e.g., through single-sideband modulation) and / or based on the RMS Parameter ratios in the amplification are adjusted. For this purpose, the ratio of the spectral centers SC of the respective speech signal segment at an onset time is determined. ten,distract and of the associated deflection signal, and the frequency of the deflection signal is adjusted so that it matches that of the speech signal as closely as possible. This can be achieved by adjusting the value of the function SC int of the interpolated spectral center point at an onset time SC int ( ten,distract ) is compared with the SC value of the selected deflection signal and a detuning parameter is determined, where positive values of the detuning parameter mean an increase in the pitch of the deflection signal by means of single-sideband modulation and negative values lead to a decrease in the pitch.
[0080] The energy (RMS) of the distraction signal is also adjusted to the energy of the speech signal segment, thus achieving a predetermined energy ratio between the distraction and speech signals. Due to their high effectiveness in reducing speech intelligibility, the distraction signals can be reproduced at a low volume, so that the overall sound pressure level at the seating positions of the listening passengers increases only minimally, thus preventing any increase in passenger annoyance or impairment and maintaining optimal local listening comfort.
[0081] In exemplary embodiments, the resulting modified spectra of the deflection signals are determined depending on a random selection of directions. ten,distractIn section 270, the spatial representation of the deflection signal is spatially variable and achieved through a binaural spatial transfer function (BRTF) by pointwise multiplication of the frequency reference points (or convolution in the time domain) of the corresponding spectra. For this purpose, a direction is randomly selected for a deflection signal in section 275. Memory 280 contains binaural spatial transfer functions (BRTFs) matching the possible directions. As already explained above for masking noise, the spatialization can be performed in the frequency or time domain. In the time domain, this is achieved through convolution with the impulse response of a selected outer ear transfer function. The spatialization of the deflection signals is preferably performed such that the deflection signals are localized as close and as present as possible to the listener's head, thus achieving a strong deflection effect. For spatial reproduction, a multi-channel (e.g.,Stereo playback is required; otherwise, single-channel playback would suffice, preferably using two speakers integrated into the neck rest.
[0082] In the case of spatialization of the deflection signal in the frequency domain, the convolution results are transformed back into the time domain by an inverse Fourier transform (IFFT) with NFFT 2 points in Section 285. The inversely transformed time blocks are then combined into a time signal in Section 290 using the overlap-add method, ensuring correct sequence and numerical accuracy. If the previous steps were already performed in the time domain, the inverse transformation and superposition of the blocks are, of course, unnecessary.
[0083] The resulting time signals are sent to the respective active neck rest of the listener. In embodiments where masking noise signals are also generated, the masking signals can be summed with the deflection signals before being output via the neck rest's loudspeakers.
[0084] The speech-signal-adapted distraction signal generates randomly spatially distributed excitation / trigger information and obscures the speech target signal improved, without significant permanently acting signal levels.
[0085] As already mentioned, the signal processing can be performed partially in the frequency domain or in the time domain. The specific values mentioned above are only examples of a possible configuration of the frequency transformation and can be modified in many ways. In one possible variation, the energy- and frequency-adapted spectra (see Section 265) are first transformed back into the time domain and then again into the frequency domain to account for spatialization—possibly with a different spectral resolution. However, it is also possible to perform the entire processing in the frequency domain. Experts in the field of digital signal processing will recognize such variations of the inventive procedure for generating speech-signal-dependent deflection signals after studying the present disclosure.
[0086] In exemplary embodiments, both masking signals—broadband masking noise and distraction signals—are summed and played back together before output. The masking noise, preferably perceived from the direction of the speaker, generates a broadband noise signal adapted to the spectral characteristics of the respective speech segment. Short distraction signals are superimposed on this noise at particularly relevant points (both temporally and frequency-wise). These distraction signals are perceived spatially near the head and lead to a particularly effective reduction in speech intelligibility, even when played back at low volume or energy. However, the combination with the broadband masking noise makes the brief switching on and off of the distraction signals less noticeable or disruptive.The overall sound pressure level at the seating positions of the listening passengers increases only minimally, and the annoyance or impairment of the passengers is not increased, or the local listening comfort is maintained as best as possible.
[0087] The above description of exemplary embodiments includes numerous details that are not essential to the invention as defined by the claims. The description of these embodiments serves to illustrate the invention and is purely illustrative, without limiting the scope of protection. It is apparent to those skilled in the art that the described elements and their technical effects can be combined in various ways, potentially leading to further embodiments covered by the claims. Furthermore, the described technical features can be used in devices and methods, for example, by programmable devices. They can be implemented, in particular, by hardware elements or by software. As is known, the implementation of digital signal processing is preferably carried out by specially designed signal processors.Communication between individual components of the described device can be wired (e.g., via a bus system) or wireless (e.g., via Bluetooth or WiFi). Protection is also expressly claimed for a computer-implemented realization and the associated program or machine code in the form of data carriers or in a downloadable format.
Claims
1. A method for masking a speech signal in a zone-based audio system (1), wherein a speech signal to be masked is received and a masking signal is generated based on said speech signal, the method comprising: acquiring said speech signal to be masked in an audio zone (I, II); characterized in that the method comprises the following method steps: transforming (105) said acquired speech signal into spectral bands; swapping (125) spectral values of at least two spectral bands; generating a masking signal adapted to said speech signal based on said swapped spectral values; and outputting said masking signal for said speech signal in another audio zone (I, II).
2. The method of claim 1, wherein generating a masking signal based on said swapped spectral values comprises: generating a wide-band noise signal; transforming (150) said generated wide-band noise signal into the frequency domain; and multiplying (155) said frequency representation of said noise signal by a frequency representation of said speech signal taking into account said swapped spectral values.
3. The method of claim 2, said frequency representation of said speech signal generated by interpolating (135) the spectral values of the bands after swapping (125) of spectral values.
4. The method of one of the previous claims, further comprising: estimating a background noise spectrum; comparing spectral values of said speech signal with the background noise spectrum; and taking into account only spectral values of said speech signal larger than the corresponding spectral values of said background noise spectrum.
5. The method of one of the previous claims, said transformation (105) of said acquired speech signal to spectral bands occurring for blocks of said speech signal and by means of a Mel-filter bank (110) and, optionally, a temporal smoothing (120) of the spectral values occurs for the Mel bands.
6. The method of one of the previous claims, presenting, when outputting in said other audio zone (I, II), said noise signal spatially by means of multi-channel playback, preferably by multiplication with binaural spectra of an acoustic transfer function.
7. The method of claim 6, said noise signal in said other audio zone (I, II) being spatially output such that it appears to originate from the dominant direction of the speaker of said speech signal to be masked.
8. The method of one of the previous claims, further comprising: determining a point of time (ti,distract) relevant for speech intelligibility in said speech signal; generating (245) a distraction signal for said determined point of time (ti,distract); and outputting said distraction signal at said determined point of time (ti,distract) as a further masking signal in said other audio zone (I, II).
9. The method of claim 8, said point of time (ti,distract) relevant for speech intelligibility determined by means of extreme values of a spectral function of said speech signal, said spectral function being determined based on an addition of optionally averaged spectral values along the frequency axis (235).
10. The method of claim 8 or 9, said point of time (ti,distract) relevant for speech intelligibility verified by means of parameters of said speech signal such as zero-crossing rate (205), short-term energy (215), and / or spectral center of gravity (220).
11. The method of one of the claims 8 to 10, said distraction signal for said determined point of time (ti,distract) selected randomly from a set of predetermined distraction signals and / or adapted to said speech signal regarding a spectral characteristic and / or its energy (265).
12. The method of one of the claims 1 to 11, presenting, when outputting by means of multi-channel playback in said other audio zone (I, II), said masking signal spatially, preferably by multiplication (270) with binaural spectra of an acoustic transfer function (280), said masking signal being spatially output in said other audio zone (I, II) such that it appears to originate from a random direction and / or in the vicinity of a listener's head in said other audio zone (I, II).
13. An apparatus (A, B) for the generation of a masking signal in a zone-based audio system (1) that receives a speech signal to be masked and generates said masking signal based on said speech signal, comprising: means (105) for transforming said acquired speech signal to spectral bands; characterized in that the apparatus comprises the following means: means (125) for swapping spectral values of at least two spectral bands; and means for generating a masking signal adapted to said speech signal based on said swapped spectral values.
14. The apparatus of claim 13, further comprising: means for determining a point of time (ti,distract) relevant for speech intelligibility in said speech signal; means (245) for generating a distraction signal for said relevant point of time (ti,distract); and means (270) for adding the noise signal and said distraction signal and for outputting the sum signal as a masking signal; and / or means for generating a multi-channel representation of said masking signal that allows for a spatial playback of said masking signal.
15. A zone-based audio system (1) with a plurality of audio zones (I, II), one audio zone (I, II) comprising at least one microphone (5) for acquiring a speech signal and another audio zone (I, II) comprising at least one loudspeaker (4), wherein said microphone (5) and loudspeakers (4) are preferable arranged in headrests (3) of seats (2) for occupants of a vehicle, and wherein said audio system (1) comprises an apparatus (A, B) for generating a masking signal according to claims 13 to 14 that obtains a speech signal from a microphone (5) of said one audio zone (I, II) and sends said masking signal to said one or more loudspeakers (4) of said other audio zone (I, II).