Audio masking of speakers

The method and device generate a masking signal by altering the spectral structure of speech and adding spatially reproduced noise and distraction signals to prevent eavesdropping in zone-based audio systems, effectively reducing speech intelligibility without disturbing overall sound levels.

EP4167228B1Active Publication Date: 2025-12-10AUDIO MOBIL ELEKTRONIK GMBH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
EP2021203247
Authority / Receiving Office
EP · EP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-10-18
Publication Date
2025-12-10
Estimated Expiration
2041-10-18

AI Technical Summary

Technical Problem

Existing communication technologies fail to effectively prevent unwanted eavesdropping in public spaces without causing unpleasant disturbances, such as increased noise levels, particularly in vehicles and public transport, where confidential conversations can be overheard due to zone-based audio systems.

Method used

A method and device for generating a masking signal in a zone-based audio system that alters the spectral structure of a speech signal by swapping spectral bands and adding a broadband noise signal, combined with spatial reproduction and distraction signals at specific speech onsets, to reduce speech intelligibility without significantly increasing overall sound levels.

Benefits of technology

Effectively prevents eavesdropping by reducing speech intelligibility in neighboring zones while maintaining minimal sound level increases and listener comfort, ensuring private conversations remain undisturbed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF0001
    Figure IMGF0001
  • Figure IMGF0002
    Figure IMGF0002
  • Figure IMGF0003
    Figure IMGF0003
Patent Text Reader

Abstract

The present disclosure relates to a method for masking a speech signal in a zone-based audio system, comprising: capturing a speech signal to be masked in one audio zone; transforming the captured speech signal into spectral bands; swapping spectral values ​​of at least two spectral bands; generating a noise signal based on the swapped spectral values; and outputting the noise signal as a masking signal for the speech signal in another audio zone.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present disclosure relates to the generation of a masking signal for speech in a zone-based audio system.

[0002] Modern communication technologies and their ever-increasing coverage enable communication to take place almost anywhere, for example, in the form of telephone conversations. In public spaces, other people can often overhear such conversations and understand their content. This is particularly problematic when the conversations are confidential, private, or business-related. Such a scenario exists on public transportation, such as trains or airplanes, but also in private vehicles, such as taxis or rented limousines. In these cases, in addition to the speaker, other people are in fixed positions, for example, in assigned seats. Often, these seats have an associated audio system or at least components thereof.For example, speakers for individual playback of audio content may be provided in these seats, for example integrated into neck rests, which is also referred to as a zone-based audio system.

[0003] Besides telephone conversations, the problem of unwanted eavesdropping can also occur in conversations between people. For example, two passengers in the back of a taxi might be discussing a confidential topic, and the driver might not want to overhear them.

[0004] It is known from current technology that unwanted eavesdropping can be reduced by playing loud noise. However, this increases the noise level for everyone involved and is perceived as an unpleasant disturbance that can also affect attention and reaction time, which is particularly undesirable in road traffic.

[0005] US Patent 2012 / 016665 A1 discloses a device for generating a masking signal, wherein a CPU analyzes the speech utterance rate of a received audio signal. The CPU then copies the received audio signal into a multitude of audio signals and performs the following processing for each of the audio signals. Specifically, the CPU divides each of the audio signals into frames based on a frame length determined by the speech utterance rate. A reversal process is performed on each of the frames to replace one waveform of the frame with an inverted waveform, and a windowing process is performed to achieve smooth transitions between the frames. Subsequently, the CPU randomly reorders the sequence of the frames and mixes the multiple audio signals to generate a masking audio signal.

[0006] In "Aircraft noise and speech intelligibility in an outdoor living space" a "Partial Loudness Model" is developed that predicts the loudness of a signal in the presence of a masking sound signal, taking into account masking across spectral bands and the effects of temporal masking of time-varying noise.

[0007] This document addresses the technical task of generating a masking signal in a zone-based audio system that reduces unwanted eavesdropping on a conversation without causing any unpleasant disturbance.

[0008] The problem is solved by the features of the independent claims. Advantageous embodiments are described in the dependent claims.

[0009] The invention is described in the attached set of claims.

[0010] According to a first aspect, a method for masking a speech signal in a zone-based audio system is disclosed. The method comprises capturing a speech signal to be masked in an audio zone, for example, by means of one or more conveniently placed microphones, which may be located, for instance, in the headrest of a seat. The speech signal may originate from the local speaker of a telephone conversation or belong to a conversation between people present. The captured speech signal is then transformed into spectral bands, which can be done, for example, using an FFT and Mel filters. Furthermore, the method involves swapping spectral values ​​of at least two spectral bands, thereby altering the spectral structure of the speech signal without changing its overall energy content. Subsequently, a (preferably broadband) noise signal is generated based on the swapped spectral values.The generated noise signal, while bearing some resemblance to the spectrum of the speech signal, does not perfectly match it, as the spectral structure of the speech signal is no longer fully preserved due to the band swapping. Such a noise signal, with a similar but not identical spectrum to the speech signal, is well-suited as a masking signal for the speech signal. It should be noted that any number of bands can be swapped (e.g., all of them), with more band swaps resulting in greater variation in the noise spectrum. Finally, the noise signal is output as a masking signal in a different audio zone with minimal energy input to make it more difficult for a person located at that listening position to overhear the conversation by reducing speech intelligibility for them.

[0011] Generating a noise signal based on swapped spectral values ​​can involve creating a broadband noise signal, for example, using a noise generator, and transforming the generated noise signal into the frequency domain. Furthermore, the frequency representation of the noise signal can be multiplied by a frequency representation of the speech signal, taking the swapped spectral values ​​into account. This frequency-domain multiplication produces a noise spectrum that essentially corresponds to that of the speech signal after the spectral bands have been swapped—that is, it is similar to, but not identical to, the speech spectrum. A similar effect can also be achieved through convolution in the time domain.

[0012] The frequency representation of the speech signal can be generated by interpolating the spectral values ​​of the bands (for example, in the mel range) after swapping the spectral values. This interpolation generates the necessary values ​​at the frequency reference points for multiplication with the noise spectrum from the (relatively few) spectral values ​​of the bands.

[0013] The method can further involve estimating a background noise spectrum (preferably at the listening position) and comparing spectral values ​​of the speech signal with the background noise spectrum. The comparison of spectral values ​​preferably (but not necessarily) takes place within the spectral bands (e.g., mel bands), which means that the background noise spectrum must also be represented in these spectral bands. Furthermore, only spectral values ​​of the speech signal that are larger than the corresponding spectral values ​​of the background noise spectrum (or bear a predetermined ratio to them) can be considered for further processing (e.g., the interpolation mentioned above). Spectral components of the speech signal that are already masked by the background noise do not need to be considered for generating the masking signal and can be suppressed (e.g., by setting them to zero).Background noise can be accounted for both before and after the swapping of spectral values. In the former case, the spectral bands being compared still match exactly, and the background noise is correctly accounted for. In the latter case, swapping bands and attenuating low-energy bands in the speech signal introduces an additional variation into the noise spectrum, which can lead to increased masking. This allows for a masking signal adapted to the background or environment, which can be output to the listener's audio field with minimal energy input.

[0014] The transformation of the captured speech signal into spectral bands can be performed for blocks of the speech signal and using a Mel filter bank. Optionally, it is possible to perform time smoothing of the spectral values ​​for the Mel bands, e.g., in the form of a moving average.

[0015] In a further embodiment of the invention, the noise signal can be spatially represented during output by means of a multi-channel (i.e., at least two-channel) playback. For this purpose, a multi-channel representation of the masking signal can be generated, enabling spatial reproduction of the masking signal. For two-channel systems, this can preferably be achieved by multiplication with binaural spectra of an acoustic transfer function. The spatial reproduction enhances the effect of the masking signal on concealing speech at the listening position, particularly when the noise signal is output spatially in the other audio zone in such a way that it appears to originate from the direction of the speaker of the speech signal to be masked.

[0016] In addition to the masking signal described above, which is based on a broadband noise signal adapted to the speech signal, a further component can be generated for the masking signal and output together to the listener in the second audio zone. For this purpose, the method can involve determining a point in the speech signal relevant to speech intelligibility (e.g., the presence of consonants) and generating a suitable distraction signal for that specific point in time. The distraction signal can then be output at that specific point in time as a further masking signal in the other audio zone, resulting in a localized additional obfuscation (masking) of the speech content during speech onsets. Since the distraction signal is only output at specific relevant points in time, it does not significantly increase the overall sound level and does not cause any significant impairment.

[0017] The relevant time for speech intelligibility can be determined based on extreme values ​​(e.g., local maxima, onsets) of a spectral function of the speech signal. This spectral function is determined by summing spectral values ​​along the frequency axis. The spectral values ​​can be smoothed beforehand in the temporal and / or frequency direction. After summing the spectral values ​​along the frequency axis, the sums can optionally be logarithmized. To generate local maxima for detecting relevant time points, the (optionally logarithmized) sums can be time-differentiated.

[0018] Furthermore, the time points relevant for speech intelligibility can be verified using parameters of the speech signal, such as zero-crossing rate, short-term energy, and / or spectral center. It is also possible to consider temporal constraints for extreme values, such as requiring them to have a predetermined minimum time interval.

[0019] The distraction signal for a specific time can then be randomly selected from a set of predefined distraction signals. These can be stored in a memory for selection. It has proven advantageous to match the distraction signal to the speech signal with respect to its spectral characteristics and / or energy. In this way, the spectral center of the distraction signal can be aligned with the spectral center of the corresponding speech segment at the specific time, for example, by means of single-sideband modulation. A speech segment with a high spectral center can thus be masked with a distraction signal that also has a high spectral center (possibly even the same spectral center), leading to greater masking effectiveness.The energy of the distraction signal can also be adjusted to match the energy of the speech segment in order to avoid creating a masking signal that is too loud and excessively disruptive.

[0020] In a further embodiment of the invention, the distraction signal can be represented during output by means of a multi-channel spatial reproduction, preferably by multiplication with binaural spectra of an acoustic transfer function, thereby generating a multi-channel (at least two-channel) representation of the distraction signal, which enables a spatial reproduction of the distraction signal. This spatial reproduction enhances the effect of the distraction signal on masking speech at the listening position, particularly when the distraction signal is output spatially in the other audio zone in such a way that it appears to originate from a random direction and / or near the listener's head in that other audio zone. This spatialization reduces the distinguishability of the speech and distraction signals, or makes it more difficult to overhear the speech signal due to the distraction signal, and thus reduces the energy required for the distraction signal.

[0021] The processing of the speech signal and the generation of a masking signal described above are preferably carried out in the digital domain. This requires steps not described in detail, such as an analog-to-digital conversion and a digital-to-analog conversion, which, however, will be self-evident to a person skilled in the art after studying the present disclosure. Furthermore, the above method can be implemented wholly or partially by means of a programmable device, which in particular includes a digital signal processor and the necessary analog-to-digital converters.

[0022] According to a further aspect of the invention, a device for generating a masking signal in a zone-based audio system is proposed, which receives a speech signal to be masked and generates the masking signal based on the speech signal. The device comprises means for transforming the acquired speech signal into spectral bands; means for interchanging spectral values ​​of at least two spectral bands; and means for generating a noise signal as a masking signal based on the interchanged spectral values.

[0023] The embodiments of the method described above can also be applied to this device. The device can further comprise: means for determining a point in time relevant to speech intelligibility within the speech signal; means for generating a distraction signal for the relevant point in time; and means for adding the noise signal and the distraction signal and outputting the sum signal as a masking signal.

[0024] In a further embodiment of the device, it also includes means for generating a multi-channel representation of the masking signal, which enables a spatial reproduction of the masking signal.

[0025] According to a further aspect of the invention, a zone-based audio system with a plurality of audio zones is disclosed, wherein at least one audio zone has a microphone for capturing a speech signal and another audio zone has at least one loudspeaker. The microphone and loudspeaker can be arranged in the headrests of seats for vehicle occupants. It is also possible for both audio zones to have a microphone and loudspeaker. The audio system includes a device, as shown above, for generating a masking signal, which receives a speech signal from a microphone of one audio zone and sends the masking signal to the loudspeaker(s) of the other audio zone.

[0026] Another aspect of the present disclosure concerns the generation of a distraction signal as a masking signal, independent of the aforementioned noise signal, as described above. A corresponding method for masking a speech signal in a zone-based audio system comprises: capturing a speech signal to be masked in one audio zone; determining a point in time relevant to speech intelligibility within the speech signal; generating a distraction signal for that specific point in time; and outputting the distraction signal at that specific point in time as a masking signal in the other audio zone. The possible embodiments of the method correspond to the embodiments described above in combination with the generated noise signal.

[0027] A corresponding device for generating a distraction signal as a masking signal in a zone-based audio system is also disclosed. This device receives a speech signal to be masked and generates the masking signal based on the speech signal. It includes means for determining a point in time relevant to speech intelligibility within the speech signal; means for generating a distraction signal for the relevant point in time; and means for outputting the distraction signal as a masking signal. Optionally, means for generating a multi-channel representation of the masking signal, enabling spatial reproduction of the masking signal, may be provided.

[0028] The features described above can be combined in many ways, even if such a combination is not explicitly mentioned. In particular, features described for a process can also be used for a corresponding device, and vice versa.

[0029] Exemplary embodiments of the invention will now be described in more detail with reference to the schematic drawing. These show: Fig. 1 schematically an example of a zone-based audio system; Fig. 2 schematically another example of a zone-based audio system; Fig. 3 schematically another example of a zone-based audio system with two zones; Fig. 4 schematically another example of a zone-based audio system with multiple zones; Fig. 5 an example of a block diagram for generating a broadband masking signal for obfuscating speech; and Fig. 6An example of a block diagram for generating a distraction signal to obfuscate speech.

[0030] The exemplary embodiments described below are not limiting and are purely illustrative. For the purpose of illustration, they include additional elements that are not essential to the invention. The scope of protection is to be determined solely by the accompanying claims.

[0031] The following examples allow vehicle occupants in any seating position to conduct undisturbed private conversations, such as phone calls with other people outside the vehicle. For this purpose, an audio masking signal is generated and transmitted to other vehicle occupants, disrupting their perception of the conversation and making it difficult, or ideally impossible, for them to overhear the private conversation. This creates a private space for the speaker, allowing them to conduct private conversations undisturbed, without the risk of other vehicle occupants overhearing confidential information. The conversation could be, for example, a telephone call or a conversation between vehicle occupants.In the latter case, there are two speakers who alternately emit speech signals that other inmates should not understand if possible, while of course ensuring that the speech intelligibility between the two participants in the conversation is not impaired.

[0032] Similar scenarios arise more generally when individuals are located in acoustic zones or environments within a space, each equipped with separate acoustic playback devices. Such acoustic zones can exist, for example, in means of transport such as vehicles, trains, buses, airplanes, ferries, etc., where passengers are seated in seats equipped with individual acoustic playback devices. However, the proposed approach to creating private acoustic zones is not limited to these examples. It can be applied more generally to situations where individuals are located in specific positions within a space (e.g., in theater or cinema seats) and can be exposed to sound through individual acoustic playback devices, and where there is a possibility of capturing the speech signals of a speaker whose speech should not be understood by the other occupants.

[0033] In one embodiment, a zone-based audio system is provided to create private acoustic zones at each passenger seat in a vehicle, or more generally, an acoustic environment. The individual components of the audio system are networked and can interact and exchange information / signals. Figure 1 Figure 1 schematically shows an example of such a zone-based audio system. A user or passenger is located in a seat 2 with a headrest 3, which has two loudspeakers 4 and two microphones 5.

[0034] Such a zone-based audio system has one, preferably at least two, loudspeakers 4 for the active acoustic reproduction of personal and individual audio signals, which should not be perceived, or only minimally perceived, by neighboring zones. The loudspeaker(s) 4 can be installed in the headrest 3, the seat 2 itself, or in the vehicle's headliner. The loudspeakers have a sufficiently optimized acoustic design and can be controlled via appropriate signal processing to minimize the acoustic impact on neighboring zones.

[0035] Furthermore, such an audio zone has the capability to record the speech of the occupant in the primary acoustic zone independently of the neighboring zones and the signals actively reproduced therein. For this purpose, one or more microphones 5 can be integrated into the seat 2 or the headrest 3, or placed in the immediate acoustic environment of the zone and the occupant, as shown in Figure 2 is shown schematically. Preferably, the microphones 5 are arranged so that they enable the best possible capture of the speech of the occupant making the phone call. If a microphone can be placed in the immediate vicinity of the speaker's mouth (like the middle microphone in Figure 2Generally, a single microphone is sufficient to capture the speaker's audio signals with adequate quality. For example, the microphone of a telephone headset can be used to record speech signals. Otherwise, two or more microphones are advantageous for capturing speech in order to record it more accurately and, above all, more precisely using digital signal processing, as explained below.

[0036] The speaker's audio zone can have appropriate signal processing to record the primary occupant's speech signals as undisturbed as possible and unaffected by neighboring zones and prevailing disturbances in the environment (wind, rolling noise, ventilation, etc.).

[0037] The voice signal of the vehicle occupant making a phone call is thus captured at the seating position (either directly by a suitably positioned microphone or indirectly by means of one or more remote microphones with appropriate signal processing) and separated from any interfering signals, such as background noise.

[0038] From this speech signal, a masking signal, also referred to as a speech obfuscation signal, can be generated for an overheard passenger. In some embodiments, a broadband masking signal adapted to the speech to be obfuscated is generated for this passenger. Additionally or alternatively, distraction signals can also be generated at the individual speech onsets within the primary speaker's speech. These are short interference signals that are emitted at specific speech segments important for speech intelligibility and can also be adapted to the speech to be obfuscated. These distraction signals are emitted with temporal overlap with the speech segments relevant for speech intelligibility in order to reduce the information content for the listener and thus improve speech intelligibility.to impair their interpretation (informational masking) without significantly increasing the overall sound level.

[0039] Adapted to the specific local acoustic requirements, these masking signals can be played back spatially (multi-channel), creating a spatial perception of the masking signals. In this way, eavesdropping from the seating positions of the listening individuals can be avoided as effectively as possible.

[0040] The proposed approach ensures that the overall sound pressure level at the seating positions of the listening passengers increases only minimally and that the annoyance or impairment (annoyance) of the passengers is not increased, or that local listening comfort is maintained as much as possible, in contrast to an approach in which a loud background noise is simply emitted to mask the speech (energetic masking).

[0041] Figure 3Figure 1 illustrates the functionality and basic system structure of an embodiment for two audio zones. The speech signals of the occupant in the primary acoustic zone I are captured by microphones 5 located in the speaker's headrest 3 of that zone and subjected to a first digital signal processing A to record the speech signals of the primary occupant as undisturbed as possible and unaffected by neighboring zones and ambient disturbances (wind, road noise, ventilation, etc.). Alternatively, the microphone(s) 5 can also be located in front of the speaker, as shown in Figure 1. Figure 2This can be depicted, for example, in the rear part of the front passenger's headrest, or in the headliner, steering wheel, or dashboard. In the example shown, the eavesdropper is in the seat directly in front of the speaker, but this is not mandatory, and the eavesdropper can be located anywhere else within the vehicle.

[0042] The processed speech signals are then fed to a second signal processing unit B, which generates appropriate speech masking signals to reduce the speech intelligibility of the eavesdropping occupant. These speech masking signals are then output via loudspeakers 4' in the second acoustic zone II. These loudspeakers are located, for example, in the headrest 3' of the eavesdropping occupant to ensure the most direct and undisturbed reproduction of the speech masking signals possible. As mentioned previously, a speech masking signal can consist of a broadband masking signal adapted to the speech signal of the primary occupant and / or a distraction signal that begins at specific points in speech. In this way, acoustic zones can be designed to be so private that unwanted eavesdropping across the boundary of an acoustic zone is significantly hindered.

[0043] An alternative solution—similar to active noise cancellation—reduces the estimated speech signals at the respective listening and microphone locations by actively applying adaptive cancellation signals. However, since the listening position is easily variable in practice, and the listening and microphone locations are often several centimeters apart, this approach can only actively reduce speech signal components up to approximately 1.5 kHz. Because speech intelligibility is primarily dominated by consonants and thus signal components with frequencies above 2 kHz, this approach alone is insufficient and, at best, problematic. If the system is not properly calibrated (e.g., incorrectly adjusted to the head position), the cancellation signals can carry precisely the relevant private information and even amplify it, thus increasing rather than decreasing speech intelligibility.In contrast, the proposed approach is less sensitive to the exact head positions of the speaker and the listener and allows for a reduction in speech intelligibility even of higher frequency speech components such as consonants.

[0044] Due to the modularity of the proposed approach, implementation examples with multiple audio zones are also conceivable, such as in mass transport (railway, airplane, train) or other fields of application (entertainment, cinema, etc.). Figure 4 This schematically illustrates such a multi-zone approach using a multi-row vehicle as an example, in which six acoustic zones are provided. As before, loudspeakers and microphones are integrated into the passengers' headrests, although the microphones can also be positioned in other locations in front of the respective speakers to optimize speech signal capture. Similar to in Figure 3In this example, it is assumed that the speaker is sitting behind the unwanted eavesdropper (here, the driver). However, the speaking occupant's voice signals can be used in the same way to generate masking or concealment signals for occupants other than the driver, and even for multiple unwanted eavesdroppers. Naturally, the speaker could also be located in a different place within the vehicle than the one described in the example. Figure 4 The example shown illustrates this approach, which can be applied generally to all scenarios where a speaker's speech is captured and the generated speech obfuscation signals can be specifically output to the unwanted eavesdropper(s).

[0045] As mentioned at the beginning, the speech signals could be a telephone conversation that the speaker is having with an external person outside the room containing the acoustic zones. Alternatively, the conversation could also be taking place between people in the room, for example, between the person in Figure 4 The system detects the speaker and the occupant to their right. In this case, the zone-based audio system must perform the same signal processing for the second speaker as for the first speaker, so that their speech is also captured and processed to generate appropriate masking signals for the listener(s). When the two speakers alternate, only the current speaker needs to be identified and the masking signals associated with that speaker output. If both speakers speak simultaneously, both masking signals can be output at the same time.

[0046] The following describes the necessary signal processing steps for an exemplary use case. In this use case, a rear-left seated passenger is conducting a telephone conversation with a person outside the vehicle as the internal speaker. In addition to the internal speaker's voice, the voice of the external speaker, for example, output from the internal speaker's headrest speaker, can also be processed. For End Speaker Signal The speech to be masked is then recorded. This speech is then retouched or masked for the driver listening from the "front left" position. Of course, this is only one possible scenario, and the proposed methods can be applied generally to all possible configurations of speaker and listener positions.

[0047] The signal estimated using digital signal processing A say yesThe base parameter for the subsequent generation of the masking or obfuscation signal is provided for the speech signal to be masked. The speech signal to be masked can be the active internal speaker in the vehicle compartment and / or the external speaker outside. The obfuscation signal can be a broadband masking signal and / or distraction signals. These generated signals ( send to: out LS-Left & LS-RightThe signals are reproduced via the active neck support at the listening position. In exemplary embodiments, both masking signals are generated, added, and reproduced together to have a stronger effect on the listener and impair their intelligibility. The combination of the two masking signals creates a synergistic effect in reducing speech intelligibility. The continuous broadband masking signal generates background noise, whereby the volume (energy) of the signal can be reduced compared to outputting only a noise signal, thus achieving a less disruptive effect. By selectively outputting the distraction signals at appropriate positions (speech onsets), the intelligibility of these speech segments (e.g.,(for consonants) are disrupted without significantly increasing the overall energy of the obfuscation signal or causing additional unpleasantness to the listener. It has even been found that the distraction signals are perceived as less unpleasant when presented together with the noise signal.

[0048] Figure 5 This shows a schematic block diagram for generating a broadband speech-signal-dependent mask. The input signal is the speech signal to be masked. say yes The resulting two-channel output signals (out LS-Left & LS-Right) are sent to the active neck support at the listening position, possibly superimposed with distraction signals, and output to the listening person via loudspeakers attached to / in the neck support.

[0049] The following describes in detail the signal processing steps for generating a broadband noise signal for speech masking according to one exemplary embodiment. It should be noted that not all steps are always necessary and some steps can be performed in a different order, as those skilled in the art will recognize. Furthermore, some calculations can be performed equivalently in the frequency domain or the time domain.

[0050] First, the speech signal say yes The signal is transformed into the frequency domain and smoothed both temporally and in the frequency direction. For this purpose, the speech signal is first processed in section 100. say yesThe signal is divided into blocks (for example, 512 samples at a sampling rate of fs = 44.1 kHz are arranged into blocks with a duration of 11.6 ms and 50% overlap). Subsequently, each signal block is transformed into the frequency domain in Section 105 using a Fourier transform with NFFT 1 = 1024 points.

[0051] In a further step (110), the Fourier spectra are filtered using a Mel filter bank with M = 24 bands—that is, the spectra are spectrally compressed by the Mel filter bank. The filter bank can consist of overlapping bands with a triangular frequency response. The center frequencies of the bands are equidistantly spaced across the Mel scale. The lowest frequency band of the filter bank starts at 0 Hz, and the highest frequency band ends at half the sampling rate (fs). For each band of the filter bank, a short-term energy value (RMS level or specific loudness response of the individual Mel bands) is calculated for each signal block in section 115 of the block diagram. These short-term energy values ​​are time-averaged over MA = 120 blocks in section 120 using a moving average (120 blocks correspond to approximately 700 ms).

[0052] In exemplary implementations described in Section 125, these dynamic loudness profiles in the immediate frequency environment are interchanged (scrambled). For this purpose, the loudness values ​​of the bands are swapped according to the following table, where the assignment of the band "in" is determined by the corresponding position in the row below, "out". For example, the loudness value of band number 2 is assigned to band number 4, and the value of band 4 is assigned to band 5, whose value is assigned to band 3, and so on. This results in swaps of loudness values ​​with adjacent or next-but-one bands; that is, the difference between a Mel band and a swapped band is, in this example, a maximum of two Mel bands. Of course, the table shown is only one possible example of band swapping, and other implementations are possible.

[0053] The proposed band swapping technique "scrambles" the loudness values, creating a degree of "disorder" in the loudness distribution for a given speech segment. This alters the description of its spectral energy, or loudness distribution, without changing the overall energy or loudness of the speech segment. For example, a particularly high energy content in one band is shifted to another, or low energy (loudness) in one band is transformed into an adjacent band. It has been shown that this redistribution of energy to neighboring bands can generate a particularly effective broadband noise signal, which reduces the intelligibility of the corresponding speech segment more significantly than without band swapping.By swapping / rotating the sequence of bins in the time-dynamic profiles of the masking bands, the transmission of speech information within the noise signal is avoided. If one were to capture the speech energy in frequency bands (e.g., mel bands as described above) and directly modulate these time-dynamic energy profiles onto a noise signal, also divided into equal frequency bands, the speech content would become audible—and more intelligible if narrow frequency bands are used. This effect is significantly reduced by swapping the loudness values ​​in the bands.

[0054] The potentially reversed dynamic loudness curves can be adjusted using the current background spectra (including all interfering noise) in section 130 of the block diagram to assess background noise and the environmental situation. For this purpose, the background noise is captured, for example, at the listening position, and the background spectra are determined by frequency transformation and temporal and frequency averaging, similar to the process used for the speech signal. Preferably, a microphone positioned at the listening position is used for this purpose. Alternatively, microphones located elsewhere (but as close as possible to the listening position) can also be used to capture the background noise at the listening position. Only those bands of the speech signal that lie above the background spectrum need to be considered when generating the masking signal.Speech bands whose energy is below that of the corresponding background noise band can be disregarded, as they are irrelevant for speech intelligibility or are already masked by the background noise. This can be achieved, for example, by setting the loudness value of such speech bands to zero. In other words, if a frequency band is already masked by strong background noise, no additional masking signal is generated in that frequency band. Thus, the system determines, based on the situation, which signal components of the broadband masking noise are used to obscure the speech.

[0055] In section 135, the resulting audibility thresholds (frequency axis sampled at 24 frequencies corresponding to the 24 center frequencies of the Mel filter bank) are interpolated at all frequency points of the Fourier transform. This interpolation generates a spectral value for the speech signal across the entire frequency range of the Fourier transform, for example, 1024 values ​​for the aforementioned Fourier transform with NFFT 1 = 1024 points.

[0056] Finally, in section 155, the frequency reference points are multiplied (or convolutioned in the time domain) of the frequency values ​​thus generated by a noise spectrum. This can be obtained by a noise generator (not shown), whose noise signal is analogous to the speech signal. say yesThe process involves block segmentation 145 and Fourier transformation 150 with identical dimensions. This generates a broadband noise signal as a masking signal with a similar frequency response (apart from the interchange and zeroing of sections 125 and 130) to the speech signal. Alternatively, the masking signal can also be generated in the time domain by convolution of the noise signal with the spectral values ​​of the speech signal processed as described above (see sections 100 to 135), which have been transformed back into the time domain. By switching between the frequency and time domains, different frequency resolutions or time durations can be used in the various processing steps. For each block of the speech signal, a broadband noise spectrum adapted to the speech segment of that block is thus generated.

[0057] In exemplary embodiments, Section 160 describes a spatial processing procedure involving point-by-point multiplication of the frequency support points (or convolution in the time domain) with binaural spectra of an acoustic transfer function that corresponds to the source direction of the speaker (or the dominant direction of the energy center of the speech signal to be masked) from the perspective of the listening person. The source direction of the speaker is known from the spatial arrangement of the acoustic zones. In the example in Figure 4 In the example shown, the speaker's source direction is directly behind the listener. In embodiments with spatial orientation of the masking signal, multi-channel playback (e.g., using two loudspeakers) is required. Otherwise, single-channel playback is sufficient, preferably also using two loudspeakers positioned in the headrest of the listener.

[0058] The broadband masking signal can thus be spatially reproduced and adapted to the target direction of the direct signal or the prominently perceived direction of the speaker. Due to the binaural loudness addition, this results in significantly improved masking with lower level excesses of the masking noise.

[0059] In Section 165, an inverse transform (IFFT) of the two spectra resulting from spatial playback (per block) into the time domain is performed, and the blocks are superimposed using the overlap-add method (see Section 170). It is noted that this results in a multi-channel signal for spatial playback, which can be played back, for example, via stereo playback. If the previous steps have already been performed in the time domain, the inverse transform and the superimposition of the blocks are, of course, unnecessary.

[0060] The resulting time signals are sent to the respective active neck rest of the listener. In embodiments where distraction signals are also generated, the masking signals can be summed with the distraction signals before being output via the neck rest's loudspeakers.

[0061] As mentioned previously, signal processing can be performed partially in the frequency domain or in the time domain. The specific values ​​mentioned above are only examples of a possible configuration and can be modified in many ways. For instance, a frequency resolution of the FFT transformation with fewer than 1024 points or a division of the Mel filters with more or fewer than 24 filters is possible. It is also possible to perform the frequency transformation of the noise signal with a different block size and / or FFT configuration than that of the speech signal. In this case, the interpolation in Section 135 would need to be adjusted accordingly to generate suitable frequency values. In another variation, the block-wise calculated masking noises are first transformed back into the time domain after interpolation and then again into the frequency domain to spatialize the sound—if necessary.with a different spectral resolution - to be taken into account. The person skilled in the art will recognize such variations of the inventive procedure for generating a broadband speech-signal-dependent masking signal after studying the present disclosure.

[0062] In exemplary implementations, short-duration deflection signals are used instead of masking noise. These signals are adapted, in terms of time and / or frequency, to sections of the speech signal that are particularly relevant for intelligibility. An example of generating such deflection signals is described below. Figure 6The diagram schematically shows an example of a block diagram for generating speech-signal-dependent distraction signals. The eavesdropper is distracted at signal-dependent, defined times. For this purpose, the critical times (ti,distract) are determined based on three information parameters in the speech signal: spectral centroid "SC" (corresponding approximately to the pitch), short-term energy "RMS" (corresponding approximately to the loudness), and number of zero crossings "ZCR" (for distinguishing speech signal from background noise).

[0063] A digital memory contains a series of pre-selected distraction signals (e.g., bird calls, chirps, etc.) with corresponding parameters (SC and RMS), determined through additional pre-analysis. Suitable distraction signals preferably have the following characteristics: Firstly, they are natural signals familiar to the listener from other situations / daily life and therefore unrelated to the signal and context to be masked. Secondly, they are characterized by being acoustically distinctive signals of short duration and exhibiting a broadband spectrum. Other examples of such signals include the sound of water dripping or rippling, or brief gusts of wind. Typically, the distraction signals are longer than the relevant speech segments (e.g., consonants) and completely mask them.It is also possible to store distraction signals of different lengths and select them to match the duration of the current critical time.

[0064] A distraction signal is selected and adjusted in time and frequency to the current speech segment. This adjusted distraction signal can then be reproduced from a virtual spatial position for the listener. For spatialization (BRTF), short impulse responses (256 points) can be used to simulate the outer ear transfer function, ensuring that these distraction signals are localized as close and prominently as possible to the listener's head, thus achieving a strong distraction effect. Multichannel playback (e.g., in stereo) is required for spatial reproduction.

[0065] The following describes in detail the signal processing steps for generating discrete, spatially distributed, short deflection signals according to one embodiment. It should be noted that not all steps are always necessary, and some steps can be performed in a different order, as those skilled in the art will recognize. Furthermore, some calculations can be performed equivalently in the frequency domain or the time domain. Some of the processing steps correspond to those for generating broadband masking signals and therefore do not need to be repeated in embodiments that use both types of signals for speech masking.

[0066] Section 200 contains the speech signal say yes divided into blocks (BlockLength = 512 Samples, fs = 44.1kHz) with a duration of 11.6 ms and 50% overlap (HopSize = 256) (see section 100).

[0067] From these blocks XBuffer n (m), where n = block index and m = time sample, the number of zero crossings (zero-crossing rate, ZCR) per signal block is determined in Section 205. This can be done using the following formula: ZCR n = 0.5 ∗ ∑ m = 1 m BlockLength − 1 sgn XBuffer n m + 1 − sgn XBuffer n m

[0068] In section 210, each signal block is subjected to a Fourier transform with NFFT 2 = 1024 points (see section 105).

[0069] From these spectra S(k,n) with k = frequency index and n = block index, two further parameters are calculated in sections 215 and 220: the short-term energy (RMS) and the spectral centroid (SC): RMS n = ∑ k S k n 2 SC n = ∑ k = 1 NFFT 2 + 1 2 k ∗ S k n ∑ k = 1 NFFT 2 + 1 2 S k n

[0070] The short-term energy (RMS) and zero-crossing rate (ZCR) curves can be further filtered using signal-dependent thresholds, and areas that do not meet these thresholds can be hidden (e.g., set to zero). The thresholds can be chosen, for example, so that a certain percentage of the signal values ​​lie above or below them.

[0071] Each spectrum is spectrally smoothed in both directions in Section 225 using a recursive time-discrete first-order filter: H(z) = Bs(z) / As(z), where Bs = 0.3 and As(z) = 1 - (Bs-1)*z -1< (= acausal, zero-phase second-order filter).

[0072] The resulting spectra are smoothed in time in section 230 using a recursive first-order discrete-time filter: H(z) = Bt(z) / At(z), where Bt = 0.3 and At(z) = 1 - (Bs-1)*z -1<.

[0073] For the detection of speech intelligibility-relevant sections (onsets) of the speech signal (onset detection), an onset detection function is first determined in Section 235. For this purpose, the spectrally and time-averaged spectra are added along the frequency axis. The resulting signal is logarithmically and time-differentiated, with negative values ​​being set to zero. Prior to the logarithm, regularization (e.g., adding a small number at each frequency point) can be performed to avoid zero values.

[0074] This onset detection function is examined for local maxima, which must be at least a predetermined number of blocks apart. The maxima found in this way can be further filtered using a signal-dependent threshold, so that only particularly pronounced maxima remain. Local maxima of the onset detection function determined in this way are candidates for perception-relevant sections of the speech signal that are to be selectively disturbed by a distraction signal.

[0075] In exemplary implementations, the maxima determined in this way by the onset detection function in Section 240 are checked for plausibility using a logic unit based on the parameters ZCR, RMS, and SC. Only if these values ​​lie within a defined range are these maxima considered relevant, critical time points. ten,distractThis can be achieved, for example, by requiring that the values ​​of RMS, SC, and / or ZCR must meet certain logical conditions at the times of the determined maxima of the onset detection function (e.g., RMS > X1; X2). <SC<X3; ZCR> X4 with predefined thresholds X1 to X4). In exemplary implementations, for example, only those maxima are considered that lie in time intervals that satisfy the aforementioned filter conditions for RMS and ZCR (i.e., not in hidden regions). The condition that ZCR and RMS must simultaneously satisfy certain threshold conditions can also be used to filter the course of SC by preserving the values ​​of SC when the threshold conditions are met and interpolating or extrapolating intermediate values, thus creating the function SC int.

[0076] At the determined times ten,distractFrom a Bouvier of N, 250 digitally stored deflection signals are randomly selected (using section 245). Additional metadata for these deflection signals, such as SC and RMS values, are stored in memory 250.

[0077] The selected deflection signal is divided into blocks in Section 255 (see above with BlockLength 2 and Hopsize = BlockLength 2 or Overlap = 0) and then Fourier-transformed using NFFT 2 points in Section 260. The parameters of this frequency transformation can be different and independent of the above procedure for the speech signal to be masked. Alternatively, the frequency representation of a deflection signal could also be stored directly in the frequency domain.

[0078] The resulting spectra can be described in Section 265 as signal-dependent. say yes at the respective time ten,distractThe gain can be adjusted based on the S-sideband parameter ratios in the frequency range (e.g., by single-sideband modulation) and / or based on the RMS parameter ratios in the gain. For this purpose, the ratio of the spectral center points S-sideband of the respective speech signal segment at an onset time is determined. ten,distract and the associated deflection signal is formed, and the frequency of the deflection signal is adjusted so that it matches that of the speech signal as closely as possible. This can be achieved by adjusting the value of the function SC int of the interpolated spectral center point at an onset time SC int ( ten,distract ) is compared with the SC value of the selected deflection signal and a detuning parameter is determined, where positive values ​​of the detuning parameter mean an increase in the pitch of the deflection signal by means of single-sideband modulation and negative values ​​lead to a decrease in the pitch.

[0079] The energy (RMS) of the distraction signal is also adjusted to the energy of the speech signal segment, thus achieving a predetermined energy ratio between the distraction and speech signals. Due to its high effectiveness in reducing speech intelligibility, the distraction signals can be reproduced at a low volume, so that the overall sound pressure level at the seating positions of the listening passengers increases only minimally, thus preventing any increase in passenger annoyance or impairment and maintaining optimal local listening comfort.

[0080] In exemplary embodiments, the resulting modified spectra of the deflection signals are determined depending on a random selection of directions. ten,distractIn section 270, the spatial representation of the time point is spatially variable using a binaural spatial transfer function (BRTF) by pointwise multiplication of the frequency reference points (or convolution in the time domain) of the corresponding spectra. For this purpose, a direction is randomly selected for a deflection signal in section 275. Binaural spatial transfer functions (BRTFs) matching the possible directions are stored in memory 280. As already explained above for the masking noise, the spatialization can be performed in the frequency or time domain. In the time domain, a convolution with the impulse response of a selected outer ear transfer function is performed. The spatialization of the deflection signals is preferably carried out so that the deflection signals are localized as close and as present as possible to the listener's head, so that they achieve a strong deflection effect. For spatial reproduction, a multi-channel (e.g.,Stereo playback is required; otherwise, single-channel playback would suffice, preferably using two speakers integrated into the neck rest.

[0081] In the case of spatialization of the deflection signal in the frequency domain, the convolution results are transformed back into the time domain by an inverse Fourier transform (IFFT) with NFFT 2 points in Section 285. The inversely transformed time blocks are superimposed in Section 290 using the overlap-and-add method. If the previous steps were already performed in the time domain, the inverse transformation and the superimposition of the blocks are, of course, unnecessary.

[0082] The resulting time signals are sent to the respective active neck support of the listener. In embodiments where masking noise signals are also generated, the masking signals can be summed with the deflection signals before output via the neck support speakers.

[0083] The speech-signal-adapted distraction signal generates randomly spatially distributed excitation / trigger information and obscures the speech target signal improved, without significant permanently acting signal levels.

[0084] As already mentioned, the signal processing can be performed partially in the frequency domain or in the time domain. The specific values ​​mentioned above are only examples of a possible configuration of the frequency transformation and can be modified in many ways. In one possible variation, the energy- and frequency-adapted spectra (see Section 265) are first transformed back into the time domain and then again into the frequency domain to account for spatialization—possibly with a different spectral resolution. Those skilled in the art will recognize such variations of the inventive procedure for generating speech-signal-dependent deflection signals after studying the present disclosure.

[0085] In exemplary embodiments, both masking signals—broadband masking noise and distraction signals—are summed and played back together before output. The masking noise, preferably perceived from the direction of the speaker, generates a broadband noise signal adapted to the spectral characteristics of the respective speech segment. Short distraction signals are superimposed on this noise at particularly relevant points (both temporally and frequency-wise). These distraction signals are perceived spatially near the head and lead to a particularly effective reduction in speech intelligibility, even when played back at low volume or energy. However, the combination with the broadband masking noise makes the brief switching on and off of the distraction signals less noticeable or disruptive.The overall sound pressure level at the seating positions of the listening passengers increases only minimally, and the annoyance or impairment of the passengers is not increased, or the local listening comfort is maintained as best as possible.

[0086] The above description of exemplary embodiments includes numerous details that are not essential to the invention as defined by the claims. The description of these embodiments serves to illustrate the invention and is purely illustrative, without limiting the scope of protection. Those skilled in the art will recognize that the described elements and their technical effects can be combined in various ways, potentially leading to further embodiments covered by the claims. Furthermore, the described technical features can be used in devices and methods, for example, by programmable devices. They can be implemented, in particular, by hardware elements or by software. As is known, the implementation of digital signal processing is preferably carried out by specially designed signal processors.Communication between individual components of the described device can be wired (e.g., via a bus system) or wireless (e.g., via Bluetooth or WiFi). An embodiment not part of the invention relates to a computer-implemented realization and the associated program or machine code in the form of data carriers or in a downloadable format.

Claims

1. Method for masking a speech signal in a zone-based audio system, comprising: detecting a speech signal to be masked in an audio zone; transforming the detected speech signal into spectral bands; interchanging spectral values of at least two spectral bands; generating a broadband noise signal based on the interchanged spectral values; generating a masking signal adapted to the speech signal based on the interchanged spectral values and the broadband noise signal; and outputting the masking signal for the speech signal in another audio zone.

2. Method of claim 1, wherein generating a masking signal comprises: transforming the generated broadband noise signal into the frequency domain; and multiplying the frequency representation of the noise signal by a frequency representation of the speech signal taking into account the interchanged spectral values.

3. Method of claim 2, wherein the frequency representation of the speech signal is generated by an interpolation of the spectral values of the bands after the interchanging of spectral values.

4. Method of one of the preceding claims, further comprising: estimating a background noise spectrum; comparing spectral values of the speech signal with the background noise spectrum; and taking into account only spectral values of the speech signal which are greater than the corresponding spectral values of the background noise spectrum.

5. Method of one of the preceding claims, wherein the transformation of the detected speech signal into spectral bands takes place for blocks of the speech signal and by means of a Mel filter bank and optionally a temporal smoothing of the spectral values for the Mel bands takes place.

6. Method of one of the preceding claims, wherein the noise signal is spatially represented in the output in the other audio zone by means of a multi-channel reproduction, preferably by multiplication by binaural spectra of an acoustic transfer function.

7. Method of claim 6, wherein the noise signal is spatially output in the other audio zone such that it appears to originate from the dominant direction of the speaker of the speech signal to be masked.

8. Method of one of the preceding claims, further comprising: determining, in the speech signal, a point in time relevant for speech intelligibility; generating a deflection signal for the determined point in time; and outputting the deflection signal at the determined point in time as a further masking signal in the other audio zone.

9. Method of claim 8, wherein the point in time relevant for speech intelligibility is determined on the basis of extreme values of a spectral function of the speech signal, wherein the spectral function is determined based on an addition of, optionally averaged, spectral values over the frequency axis.

10. Method of claim 8 or 9, wherein the point in time relevant for speech intelligibility is verified on the basis of parameters of the speech signal, such as zero crossing rate, short-term energy and / or spectral centroid.

11. Method of one of claims 8 to 10, wherein the deflection signal for the determined point in time is randomly selected from a set of predetermined deflection signals and is adapted to the speech signal with regard to a spectral characteristic and / or its energy.

12. Method of one of claims 8 to 11, wherein the deflection signal is spatially represented in the output by means of a multi-channel reproduction, preferably by multiplication by binaural spectra of an acoustic transfer function.

13. Method of claim 12, wherein the deflection signal is spatially output in the other audio zone such that it appears to originate from a random direction and / or close to the head of a listener in the other audio zone.

14. Device for generating a masking signal in a zone-based audio system, which receives a speech signal to be masked and generates the masking signal based on the speech signal, comprising: means for transforming the detected speech signal into spectral bands; means for interchanging spectral values of at least two spectral bands; means for generating a broadband noise signal based on the interchanged spectral values; and means for generating a masking signal adapted to the speech signal based on the interchanged spectral values and the broadband noise signal.

15. Device of claim 14, further comprising: means for determining a point in time relevant for speech intelligibility in the speech signal; means for generating a deflection signal for the relevant point in time; and means for adding the masking signal and the deflection signal and for outputting the sum signal as a masking signal.

16. Device of claim 14 or 15, further comprising: means for generating a multi-channel representation of the masking signal which enables a spatial reproduction of the masking signal.

17. Zone-based audio system with a plurality of audio zones, wherein one audio zone comprises at least one microphone for detecting a speech signal and another audio zone comprises at least one loudspeaker, wherein microphone and loudspeaker are preferably arranged in neck supports of seats for occupants of a vehicle, wherein the audio system comprises a device for generating a masking signal according to claims 14 to 16, which receives a speech signal from a microphone of the one audio zone and transmits the masking signal to the loudspeaker or loudspeakers of the other audio zone.

Citation Information

Patent Citations

  • Speech signal processing apparatus for cutting out a speech signal from a noisy speech signal BACKGROUND OF THE INVENTION 1. field of invention

    DE69130687T2

  • Directional sound masking

    EP2877991A2

  • Directional sound masking

    EP2877991B1

  • Sound masking system and masking sound generation method

    US20120016665A1