Audio masking of speech

JP2024542967A5Pending Publication Date: 2025-10-10AUDIO MOBIL ELEKTRONIK GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024524500
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-17
Filing Date
2022-10-18
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing zone-based audio systems fail to effectively reduce unwanted overhearing of speech without causing unpleasant disturbances, such as increased noise levels, which is particularly problematic in public transportation and private vehicles.

Method used

A method and apparatus for generating a masking signal in zone-based audio systems that involves detecting speech signals, transforming them into spectral bands, rectifying these bands to create a noise signal with a similar but not identical spectrum, and outputting this noise signal in other audio zones to reduce speech intelligibility, while maintaining low energy input and avoiding significant noise increase.

Benefits of technology

The solution effectively reduces speech intelligibility in unintended listening areas without significantly increasing overall sound levels, thus preserving privacy and avoiding passenger discomfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To generate a masking signal in a zone-based audio system that reduces undesirable overhearing of conversation while at the same time not presenting unpleasant interference. [Solution] A method for masking a speech signal in a zone-based audio system includes detecting a speech signal to be masked in an audio zone, converting the detected speech signal into spectral bands, rectifying spectral values ​​of at least two spectral bands, generating a noise signal based on the rectified spectral values, and outputting the noise signal as a masking signal for the speech signal in another audio zone.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to generating masking signals for speech in zone-based audio systems. [Background technology]

[0002] The ever-increasing range of communication means of the prior art and their coverage allows communication almost everywhere, for example in the form of telephone calls. In public places, other people can often overhear such calls and understand their contents. This is particularly problematic for sensitive private or business calls. Such situations can occur in public transport such as trains and planes, but also in private vehicles such as taxis and rented limousines. In these cases, in addition to the speaker, for example, there are other people in the assigned seat. Such seats often have an associated audio system or at least parts thereof. For example, these seats can be equipped with loudspeakers, for example integrated in the headrests, for individual reproduction of audio content, also called zone-based audio systems.

[0003] In addition to telephone conversations, the problem of unwanted overhearing can also occur in person-to-person conversations, for example when two passengers in the back seat of a taxi are discussing sensitive matters that they do not want to be overheard by the driver. Summary of the Invention [Problem to be solved by the invention]

[0004] It is known in the prior art that unwanted overhearing can be reduced by playing back a loud noise, but this increases the noise level for all involved, which is perceived as an unpleasant disturbance and may also affect attention and reaction ability, which is particularly undesirable in road traffic.

[0005] A technical object of the present invention is to generate a masking signal in a zone-based audio system that reduces unwanted overhearing of speech while at the same time not presenting objectionable interference. This object will be solved by the features of the independent claims. Advantageous embodiments are set forth in the dependent claims. [Means for solving the problem]

[0006] According to a first aspect, a method for masking a speech signal in a zone-based audio system is disclosed. The method comprises detecting the speech signal to be masked in the audio zone by one or more appropriately positioned microphones, which may be placed, for example, in the headrest of the seat. The speech signal may originate from a local speaker of a telephone conversation or belong to a conversation between people present. The detected speech signal is then transformed into spectral bands, which may be performed, for example, using FFT and Mel filters. The method also involves rectifying the spectral values ​​of at least two spectral bands, which modifies the spectral structure of the speech signal without modifying its overall energy content. A noise signal (as wide as possible band) is then generated based on the rectified spectral values. The generated noise signal shows a certain similarity to the spectrum of the speech signal, but does not match perfectly, since the spectral structure of the speech signal is no longer fully preserved by the rectification of the bands. Such a noise signal, having a spectrum similar but not identical to the speech signal, is well suited as a masking signal for the speech signal. It should also be noted that any number of bands (e.g., all of them) can be rectified, with more bands rectified resulting in greater variation in the noise spectrum. Finally, the noise signal is output as a masking signal with the lowest possible energy input in other audio zones, making it more difficult to overhear the conversation by reducing speech intelligibility for those in listening positions.

[0007] Generating a noise signal based on the rectified spectral values ​​may involve generating a wideband noise signal using, for example, a noise generator, and transforming the generated noise signal into the frequency domain. Furthermore, a multiplication of a frequency representation of the noise signal with a frequency representation of the speech signal may be performed while taking into account the rectified spectral values. The multiplication in the frequency domain generates a noise spectrum that essentially corresponds to the noise spectrum of the speech signal after the spectral bands have been rectified, i.e., similar but not identical to the speech spectrum. A similar effect may also be achieved by a convolution in the time domain.

[0008] A frequency representation of the speech signal can be generated by interpolating the spectral values ​​of the bands (e.g., to be in the mel range) after rectifying the spectral values. Interpolating from the (relatively few) spectral values ​​of the bands generates the required values ​​at the frequency support for multiplication with the noise spectrum.

[0009] The method further comprises estimating the background noise spectrum (preferably at the listening position) and comparing the spectral values ​​of the speech signal with the background noise spectrum. The comparison of the spectral values ​​is preferably (but not necessarily) performed within a spectral band (e.g., Mel-band), which means that the background noise spectrum must also be displayed in the spectral band. Furthermore, only the spectral values ​​of the speech signal that are greater (or are a certain ratio) than the corresponding spectral values ​​of the background noise spectrum can be considered for the further procedure (e.g., the above-mentioned interpolation). Spectral components of the speech signal that are already masked by the background noise do not have to be considered for the generation of the masking signal and can be masked (e.g., by setting them to zero). The consideration of the background noise can be done both before and after the rectification of the spectral values. In the former case, the compared spectral bands still match exactly and the background noise is taken correctly. In the latter case, the rectification of the bands and the masking of the low-energy bands in the speech signal can bring additional variations in the noise spectrum, which can result in an increase in masking. This allows a masking signal that is adapted to the background or environment and can be output with the lowest possible energy input in the hearing person's speech zone.

[0010] The conversion of the captured speech signal into spectral bands can be performed on blocks of the speech signal using a Mel filter bank. Optionally, a temporal smoothing of the spectral values ​​of the Mel bands can be performed, for example in the form of a floating average.

[0011] In another embodiment of the invention, a multi-channel (i.e. at least two-channel) reproduction can be used to spatially represent the noise signal at the output. For this purpose, a multi-channel representation of the masking signal can be generated that allows spatial reproduction of the masking signal. In the case of a two-channel system, this can be preferably performed by multiplication by the binaural spectrum of the acoustic transfer function. Spatial reproduction increases the effect of the masking signal obscuring the speech at the listening position, especially when the noise signals in other audio zones are output in space such that they appear to emanate from the direction of the speaker of the speech signal to be masked.

[0012] In addition to the above masking signal based on a wideband noise signal adapted to the speech signal, another component can be generated for the masking signal, which is output together to the hearer in the second audio zone. For this purpose, the method can include determining a time point in the speech signal that is related to speech intelligibility (e.g., the presence of a consonant in the speech signal) and generating an appropriate disturbing signal for the particular time point. Then, the output of the disturbing signal at the particular time can occur as another masking signal in the other audio zone, thus providing a selective additional obfuscation (masking) of the conversational content at the speech onset. Since the disturbing signal is only emitted at the particular relevant time point, it does not significantly increase the overall sound level and does not cause significant disturbance.

[0013] The time points relevant for speech intelligibility can be determined using extrema (e.g., local maxima, onsets) of a spectral function of the speech signal, which is determined based on summing the spectral values ​​over a frequency axis. The spectral values ​​can be pre-smoothed in time and / or frequency direction. After summing the spectral values ​​over the frequency axis, the sum value can be optionally logarithmicized. The sum value, optionally logarithmicized, can be time-differentiated to generate a local maximum for detection of the relevant time points.

[0014] Furthermore, the time points relevant to speech intelligibility can be verified using parameters of the speech signal such as zero-crossing rate, short-time energy, and / or spectral centroid, and can also be subject to constraints on extreme values, such as requiring a predefined minimum time span.

[0015] The disturbing signal for a particular time can then be selected randomly from a set of predefined disturbing signals. These can be kept in a memory ready for selection. It has been found to be advantageous if the disturbing signal is adapted to the speech signal in terms of its spectral characteristics and / or its energy. In this way, the spectral centroid of the disturbing signal can be adapted to the spectral centroid of the corresponding speech segment at a particular time, for example by single-sideband modulation. Thus, a speech segment with a high spectral centroid can be masked using a disturbing signal with an equally high spectral centroid (possibly even the same spectral centroid), which results in a higher masking effect. In order not to generate a masking signal that is too large and excessively disturbing, the energy of the disturbing signal can also be adapted to the energy of the speech segment.

[0016] In another embodiment of the present invention, the disturbing signal can be represented at the output using multi-channel spatial reconstruction, preferably using multiplication by the binaural spectrum of the acoustic transfer function, thereby generating a multi-channel (at least two-channel) representation of the disturbing signal that allows spatial reconstruction of the disturbing signal. Spatial reconstruction increases the effectiveness of the disturbing signal for speech obfuscation at the co-listening position, especially when the disturbing signal in the other sound zone is spatially output in such a way that it appears to originate from random directions in the other sound zone and / or from near the listener's head. This spatialization can reduce the discernibility of the speech and the disturbing signal, or make it more difficult for the disturbing signal to overhear the speech signal, thus reducing the energy of the disturbing signal.

[0017] The speech signal processing and masking signal generation described above are preferably performed in the digital domain. For this purpose, steps are required that are not described in detail herein, such as analog-to-digital and digital-to-analog conversion, but will be apparent to those skilled in the art after reviewing this disclosure. Furthermore, all or part of the above-described methods can be realized using programmable devices, in particular those comprising digital signal processing devices and analog-to-digital converters, as required.

[0018] According to another aspect of the present invention, an apparatus for generating a masking signal in a zone-based audio system is proposed, which receives a speech signal to be masked and generates a masking signal based on the speech signal, the apparatus comprising: means for converting the detected speech signal into spectral bands, means for rectifying spectral values ​​from at least two spectral bands, and means for generating a noise signal as the masking signal based on the rectified spectral values.

[0019] The above embodiments of the method as described therein can also be applied to this apparatus, which can thus further comprise means for determining a time point in the speech signal that is related to speech intelligibility, means for generating a disturbance signal for the related time point, and means for adding the noise signal and the disturbance signal and outputting the sum signal as a masking signal.

[0020] In another embodiment of the device, the device also comprises means for generating a multi-channel representation of the masking signal, allowing spatial reproduction of the masking signal.

[0021] According to yet another aspect of the present invention, a zone-based audio system is disclosed having a plurality of audio zones, at least one of which comprises a microphone for detecting a voice signal and the other audio zones comprise at least one loudspeaker. The microphone and the loudspeaker may be located in a headrest of a passenger seat of the vehicle. It is also possible that both audio zones comprise a microphone and a loudspeaker. The audio system comprises an apparatus for generating a masking signal as described above, which receives a speech signal from a microphone of said one audio zone and transmits the masking signal to one or more loudspeakers of the other audio zone.

[0022] Yet another aspect of the present disclosure relates to generating a disturbing signal as a masking signal independent of the noise signal as described above. A suitable method for masking a speech signal in a zone-based audio system includes detecting a speech signal to be masked in one audio zone, determining a time point in the speech signal related to speech intelligibility, generating a disturbing signal for the determined time point, which may be adapted to the speech signal in terms of its spectral characteristics and / or its energy, and outputting the disturbing signal at the determined time point as a masking signal in the other audio zone. A possible embodiment of the method corresponds to the above embodiment in combination with a generated noise signal.

[0023] A suitable apparatus for generating a disturbing signal as a masking signal in a zone-based audio system is also disclosed, which receives a speech signal to be masked and generates the masking signal based on the speech signal. The apparatus comprises means for determining a time point in the speech signal related to speech intelligibility, means for generating a disturbing signal for the related time point, which may be adapted to the speech signal in terms of its spectral characteristics and / or its energy, and means for outputting the disturbing signal as the masking signal. Optionally, means for generating a multi-channel representation of the masking signal may be provided, allowing spatial reproduction of the masking signal.

[0024] The above features can be combined with one another in many ways, even if such a combination is not specifically mentioned. In particular, features described for a method can also be used in the associated device and vice versa. [Brief description of the drawings]

[0025] In the following, embodiments of the present invention will be described in detail with reference to schematic drawings. [Figure 1] FIG. 1 is a schematic diagram of an example of a zone-based audio system. [Diagram 2] FIG. 2 is a schematic diagram of another example of a zone-based audio system. [Diagram 3] FIG. 2 is a schematic diagram of another example of a zone-based audio system having two zones. [Figure 4] FIG. 2 is a schematic diagram of another example of a zone-based audio system having several zones. [Diagram 5] FIG. 1 illustrates an example block diagram for generating a broadband masking signal for speech obfuscation. [Figure 6] FIG. 1 is an example of a block diagram for generating a disturbance signal for speech obfuscation. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0026] The following embodiments are not limiting and are purely exemplary. For illustrative purposes, the following embodiments include additional elements that are not essential to the invention. The scope of the invention is defined only by the scope of the appended claims.

[0027] The following embodiment allows a passenger of a vehicle at any seat position to have an undisturbed private conversation, such as a phone call, with other people outside the vehicle. For this purpose, a sound masking signal is generated and provided to the passengers of the other vehicles, which prevents the other passengers from listening to the conversation and makes it more difficult, and at best impossible, to hear the private conversation. In this way, the privacy of the speaker is created, who can also have an undisturbed private conversation without the risk that passengers of other vehicles can eavesdrop on sensitive information. The conversation can be, for example, a phone call or a conversation between passengers of the vehicle. In the latter case, the two speakers alternately emit speech signals that are incomprehensible to other passengers, but of course the speech intelligibility between the two conversation participants should not be impaired.

[0028] A similar situation generally occurs when several people are located in an acoustic zone or acoustic environment of a room, each of which is provided with sound by a separate sound reproduction device. For example, such an acoustic zone may be present in a means of transport such as a car, train, bus, plane, ferry, etc., where passengers are located in seats, each of which is equipped with a sound reproduction device. However, the proposed approach to create private acoustic zones is not limited to these examples. It may be applied more generally to situations where several people are located in different parts of a room (e.g., seats in a theater or cinema), may be exposed to sound by separate sound reproduction means, and may capture the speech signal of a speaker whose speech is not intended to be understood by others.

[0029] In one embodiment, a zone-based audio system is provided to create private acoustic zones for each passenger seat in a vehicle, or more generally in the acoustic environment. The individual components of the audio system are interconnected and can interactively exchange information / signals. Figure 1 shows a schematic example of such a zone-based audio system 1. A user or passenger sits in a seat 2 equipped with a headrest 3, which has two loudspeakers 4 and two microphones 5.

[0030] Such a zone-based sound system comprises one (preferably at least two) loudspeakers 4 for active acoustic reproduction of personal and individual sound signals, which must not be or only slightly perceived by adjacent zones. The loudspeakers 4 can be mounted in the headrests 3, in the seats 2 themselves, or in the headliner of the vehicle. The loudspeakers have a suitable acoustic design and can be controlled via suitable signal processing to minimize the acoustic impact on adjacent zones.

[0031] Moreover, such an audio zone has the ability to record the speech of passengers in the first acoustic zone independently of adjacent zones and the signals actively played therein. For this purpose, one or more microphones 5 can be integrated into the seat 2 or headrest 3 or mounted in the zone and in the direct acoustic environment of the passenger, as shown diagrammatically in FIG. 2. Preferably, the microphones 5 are positioned so that they can detect as much as possible the speech of the passenger using the phone. If it is possible to position the microphone very close to the mouth of the person speaking (such as the central microphone in FIG. 2), a single microphone is generally sufficient to capture the voice signal of the person speaking with sufficient quality. For example, the microphone of a telephone headset can be used to capture the speech signal. Otherwise, two or more microphones are advantageous in order to capture the sound so that the speech is recorded more effectively and in particular in a more targeted manner using digital signal processing, as will be explained below.

[0032] The speaker's audio zone may have appropriate signal processing to record the main passenger's voice signal with as little disturbance as possible and without being affected by dominant disturbances in adjacent zones and the environment (wind, rolling noise, ventilation, etc.).

[0033] Thus, the voice signal of a vehicle passenger on a call is recorded at the seat position (either directly by a microphone positioned accordingly, or indirectly by one or more remote microphones with appropriate signal processing) and separated from interfering signals such as background noise.

[0034] From this speech signal, a masking signal, also referred to as speech obfuscating signal in the following, can be generated for the overhearing passenger. In an example embodiment, a wideband masking signal adapted to the speech to be obfuscated is generated for this passenger. Additionally or alternatively, disturbing signals can also be generated at individual speech onsets within the speech of the main speaker. These are short interfering signals emitted at specific speech segments that are important for speech intelligibility and can also be adapted to the speech to be obfuscated. These disturbing signals are emitted so as to overlap speech segments relevant for speech intelligibility, reducing the information content for the listener and impairing the intelligibility of the speech or its interpretation (information masking) without significantly increasing the overall sound level.

[0035] These obfuscated signals can be adapted to the respective local acoustic requirements and delivered in a spatial manner (multi-channel) such that a spatial perception of the obfuscated signals is created, thus avoiding as much as possible overhearing at the listener's seating position.

[0036] Using the approach proposed above, in contrast to the approach of simply outputting a loud noise to cover the speech (energy masking), the overall sound pressure level at the listening passengers' seats is only minimally increased and passenger discomfort is not increased or local listening comfort is maintained in the best possible way.

[0037] Figure 3 is an example of the functionality and basic system structure of an example embodiment for two audio zones. The passenger speech signal in a first acoustic zone I is recorded by a microphone 5 of this zone placed in the speaker's headrest 3 and is subjected to a first digital signal processing A in order to record the main passenger speech signal as free as possible from interference and unaffected by dominant disturbances in adjacent zones and the environment (wind, rolling noise, ventilation, etc.). Alternatively, the microphone 5 can be placed in front of the speaker, for example at the rear of the front passenger's headrest, or in the headliner, steering wheel or dashboard, as shown in Figure 2. In the illustrated example, the listener is seated in the seat directly in front of the speaker, but this does not necessarily have to be the case and the listener can be located anywhere else in the vehicle.

[0038] The speech signal thus processed is then fed to a second signal processing B, which generates a suitable speech obfuscated signal so that the speech intelligibility of the hearing passenger is reduced. The speech obfuscated signal is then output via loudspeakers 4' of the second acoustic zone II. These are arranged, for example, in the headrests 3' of the hearing passengers so as to achieve the most direct and unobstructed reproduction of the speech obfuscated signal possible. As already indicated, the speech obfuscated signal can comprise a broadband masking signal adapted to the speech signal of the main passenger and / or a disturbing signal starting at the individual speech onset. In this way, the acoustic zone can be made private, such that unwanted overhearing beyond the boundaries of the acoustic zone is significantly more difficult.

[0039] In an alternative approach similar to active noise suppression, the estimated speech signal at each listening or microphone position is reduced by actively adding an adapted clearing signal.

[0040] However, since the listening position is in practice slightly variable and the listening position and microphone position are several centimeters apart, only speech signal components up to about 1.5 kHz can be actively reduced. However, this approach alone is insufficient, and should be considered important even in the best of circumstances, since speech intelligibility is dominated mainly by consonants and thus by signal components with frequencies above 2 kHz, because in the case of improper tuning (e.g., incorrect head position adjustment), the clearing signal may accurately convey and even amplify relevant private information, thereby enhancing rather than reducing speech intelligibility. In contrast, the disclosed approach allows for low sensitivity to the exact head position of the speaker and hearer and for speech intelligibility to be reduced even in the case of higher frequency speech components such as consonants.

[0041] Due to the modularity of the disclosed approach, example embodiments involving multiple audio zones are also conceivable, for example, in mass transit (railroads, planes, trains) or other applications (entertainment, cinema, etc.). FIG. 4 shows a schematic of such a multi-zone approach using a multi-row vehicle in which six acoustic zones are provided. As before, loudspeakers and microphones are integrated in the passenger headrests, but the microphones can also be placed in other positions in front of each speaker to provide a preferred position for capturing the speech signal. As in FIG. 3, this example assumes that the speaker is sitting behind the unwanted overhearing passenger (in this case the driver). However, the voice signal of the speaking passenger can be used in the same way to generate a masking or obfuscating signal for passengers other than the driver and for several unwanted overhearers. Of course, the speaker can be in a different position in the vehicle than in the example shown in FIG. 4. The approach disclosed herein can be generally applied to all situations in which the speaker's speech can be detected and the generated speech obfuscating signal can be outputted towards one or more unwanted overhearers.

[0042] As explained at the beginning, the voice signal may be a telephone conversation that a speaker has with an external person outside the room in which the acoustic zone is located. Alternatively, the conversation may be between people in the room, for example between the speaker shown in FIG. 4 and a passenger to his right. In this case, the same signal processing as for the speaker shown must be provided for the second speaker in the zone-based voice system, so that the speech of the second speaker is also detected and processed to generate appropriate obfuscated signals for the hearer or hearers. If two speakers speak alternately, it is necessary to determine only the current speaker and output the obfuscated signal associated with this speaker. If both speakers speak simultaneously, both obfuscated signals can also be output simultaneously.

[0043] In the following, the necessary signal processing steps are described in an exemplary application, where a vehicle passenger sitting in the left rear seat is making a phone call to a person outside the vehicle as an in-vehicle talker. In addition to the in-vehicle talker's speech, the outside talker's speech (far-end talker signal), e.g. emanating from a loudspeaker in the headrest of the in-vehicle talker, can also be recorded as speech to be obfuscated. This is modified or obfuscated for a driver hearing in the "left front" position. Of course, this is only one possible situation, and the proposed procedure can be used in general for any possible configuration of talker positions and listening position arrangements.

[0044] The signal sig estimated by digital signal processing A for the speech signal to be obfuscated estprovides the basic parameters for the generation of the subsequent masking or obfuscating signal. The speech signal to be masked can be an active in-vehicle speaker and / or an external speaker in the vehicle. The obfuscating signal can be a wideband masking signal and / or a disturbing signal. These generated signals (destination: LS left and LS right) are played through the active headrest at the listening position. In the example embodiment, both obfuscating signals are generated, added and played together to give an amplified effect on overhearing and affect its intelligibility. The combination of the two obfuscating signals creates a synergistic effect of these signals in reducing speech intelligibility. The continuous wideband masking signal generates background noise, thereby reducing the amount (energy) of the signal compared to the output of only a single noise signal, and a less disruptive effect is achieved. By outputting the disturbing signals on time and at the right position (speech onset), the speech intelligibility of these speech segments (e.g., for consonants) is disrupted in a targeted manner without significantly increasing the overall energy of the obfuscated signal or causing additional unpleasant effects to the listener. It has even been found that the disturbing signals are perceived as less unpleasant when presented together with a noise signal.

[0045] FIG. 5 is a schematic block diagram for generating wideband speech signal dependent masking. The input signal is the speech signal sig to be masked. est The resulting two-channel output signal (Out: LS Left and LS Right) is sent to the active neck rest at the hearing location, superimposed with a disturbance signal if necessary, and output to the hearer by a loudspeaker mounted at or in the neck rest.

[0046] In the following, the signal processing steps for generating a broadband noise signal for speech masking according to an example embodiment are described in detail. It should be noted that, as known to those skilled in the art of digital signal processing, not all steps need to be performed, some steps may be performed in a different order, and some calculations may be equally performed in the frequency domain or the time domain.

[0047] First, the speech signal sig est To this end, first, in section 100, the speech signal sig est into multiple blocks (e.g., 512 samples with a sampling rate of fs=44.1 kHz are arranged into multiple blocks of 11.6 ms duration and 50% overlap). Then, in section 105, each signal block is transformed into the frequency domain using a NFFT1=1024-point Fourier transform.

[0048] Further, in step 110, the Fourier spectrum is filtered with a Mel filter bank of M=24 bands, i.e. the spectrum is spectrally compressed by the Mel filter bank. The filter bank may consist of overlapping bands with a triangular frequency response. The center frequencies of the bands are equally spaced across the Mel scale. The lowest frequency band of the filter bank starts at 0 Hz and the highest frequency band ends at half the sampling rate (fs). Short-term energy values ​​(RMS levels or specific loudness curves of the individual Mel bands) are calculated for each signal block of all bands of the filter bank in section 115 of the block diagram. These short-term energy values ​​are averaged over time over MA=120 blocks in the form of a sliding average (moving average, 120 blocks corresponding to approximately 700 ms) in section 120.

[0049] In the example embodiment, in section 125, these dynamic loudness curves are rectified (scrambled) in the instantaneous frequency environment. For this purpose, the loudness values ​​of the bands are rectified according to the following table, where the assignment of the bands "in" is taken from the corresponding position of the row "out" below. For example, the loudness value of band number 2 is assigned to band number 4. Band number 4 and band value 4 are assigned to band 5, the value of band 5 is assigned to band 3, etc. As a result, the loudness values ​​are rectified in the adjacent or next band, i.e. the difference between the mel band and the rectified band is at most 2 mel bands in this example. Of course, the shown table is only one possible example of how the bands can be rectified, other realizations are possible.

[0050] [Table 1]

[0051] By the proposed band rectification, loudness values ​​are "scrambled" in such a way that a certain "disorder" occurs in the distribution of loudness values ​​of the relevant speech segments, thereby modifying the description of the spectral energy or its loudness distribution without modifying the overall energy or loudness of the speech segments. For example, a particularly prominent energy content of one band is shifted to another band, or a low energy (loudness) of one band is transferred to an adjacent band. By redistributing the energy to adjacent bands, a particularly effective wideband noise signal can be generated, which has been shown to reduce the intelligibility of the relevant speech segments compared to the case without band rectification. By rectifying / reversing the sequence of bins of the time-dynamic progression of the masking bands, the transmission of speech information in the noise signal is avoided. If the speech energy is captured in frequency bands (e.g. the mel-bands mentioned above) and the amplitude of these temporal energy curves is directly modulated to the noise signal and also divided into equal frequency bands, the speech content becomes audible, and is even more intelligible if narrow frequency bands are used. This effect is greatly reduced by band rectification of the loudness values.

[0052] The dynamic loudness curve, rectified if necessary, can be adjusted in section 130 of the block diagram using the current background spectrum (including all background noise) to evaluate the background noise and the surrounding situation. For this purpose, the background noise is detected, for example, at the monitoring position, and the background spectrum is determined using frequency conversion and time and frequency averaging, as with the speech signal. Preferably, a microphone placed at the listening position is used for this purpose. Alternatively, a microphone placed elsewhere, but preferably close to the monitoring position, can be used to capture the background noise at the monitoring position. When generating the masking signal, only bands of the speech signal that lie above the background spectrum need to be considered. Speech bands whose energy is lower than the energy of the corresponding background noise bands can be ignored, since they play no role in speech intelligibility or are already masked by the background noise. This can be done, for example, by setting the loudness value of such speech bands to 0. In other words, if a frequency band is already masked by a strong background noise, no additional masking signal is generated in this frequency band. Thus, the decision as to which signal components of the broadband masking noise are used to obfuscate speech is made on a case-by-case basis.

[0053] In section 135, an interpolation of the resulting common hearing threshold (frequency axis sampled at 24 frequencies corresponding to the 24 center frequencies of the Mel filter bank) is performed at every frequency sampling point of the Fourier transform. The interpolation produces spectral values ​​of the speech signal in the full frequency range of the Fourier transform, e.g. 1024 values ​​for the above Fourier transform with NFFT1=1024 points.

[0054] Finally, in section 155, a point-by-point multiplication of the frequency grid points (or a time domain convolution) of the frequency values ​​thus generated is performed with a noise spectrum. This can be obtained by a noise generator (not shown), which noise signal is passed in the same manner as the speech signal and of the same dimensions in a block division 145 and a Fourier transform 150. In this way, a wideband noise signal is generated as a masking signal having similar frequency characteristics as the speech signal (except for the rectification and zeroing in sections 125 and 130). Alternatively, the masking signal can be generated in the time domain by convolving the noise signal with the spectral values ​​of the speech signal processed as above (see sections 100-135) and transformed into the time domain. By switching between the frequency domain and the time domain, different frequency resolutions or durations can be used in the various processing steps. Alternatively, it is also possible to perform the entire processing in the frequency domain. In this way, a wideband noise spectrum adapted to the speech segments of the block is generated for each block of the speech signal.

[0055] In an example embodiment, section 160 is followed by spatial processing using point-wise multiplication of the frequency grid points (or time domain convolution, see above) with the binaural spectrum of the acoustic transfer function corresponding to the talker's source direction as seen by the hearer (or the main direction of the energy centroid of the speech signal to be masked). The talker's source direction is known from the spatial arrangement of the acoustic zones. In the example shown in FIG. 4, the talker's source direction is directly behind the hearer. In an example embodiment with spatial directionality of the masking signal, multi-channel reproduction (e.g. using two loudspeakers) is necessary. Otherwise, single-channel reproduction is sufficient, which is preferably done by two loudspeakers placed in the neckrest of the hearer.

[0056] Thus, the broadband masking signal can be spatially reproduced and adapted to the target direction of the direct signal or the predominantly perceived direction of the speaker. The addition of binaural loudness results in a significantly improved masking and a lower level of excess of the masking noise.

[0057] In section 165, the two resulting spectra (for spatial reconstruction) (block by block) are inverse transformed (IFFT) to the time domain, and overlap of blocks is performed using overlap-and-add method (see section 170). It should be noted that in spatial reconstruction, a multi-channel signal is generated, which is reproduced, for example, by stereo reconstruction. It is understood that in the case that the previous steps have already been performed in the time domain, the inverse transformation and overlap of blocks are omitted.

[0058] The resulting time signal is transmitted to each active neck rest of the hearer, so in example embodiments where a disturbance signal is also generated, the masking signal can be summed with the disturbance signal before being output via the neck rest speaker.

[0059] As already indicated, the signal processing can be partly in the frequency domain or in the time domain, but it is quite possible to perform the entire processing in the frequency domain. The above specific values ​​are only one example of a possible configuration and can be modified in many ways. For example, a frequency resolution of the FFT transformation less than 1024 points or a division of the Mel filter by more or less than 24 filters is possible. The frequency transformation of the noise signal can also be performed with a different configuration to the block size and / or FFT of the speech signal. In this case, the interpolation of section 135 must be adjusted accordingly to generate the appropriate frequency values. In yet another variant, the masking noise calculated block by block is first retransformed into the time domain after interpolation and then back to the frequency domain to allow spatialization, possibly with a different spectral resolution. Those skilled in the art will recognize such variants of the procedure according to the invention for generating a masking signal dependent on a wideband speech signal after considering the present disclosure. In an example embodiment, instead of masking the noise, a short-term disturbing signal is used, which is adapted in terms of time and / or frequency to sections of the speech signal that are particularly relevant for intelligibility. By way of example, the generation of such a disturbing signal is described below. Figure 6 shows, in a simplified manner, an example of a block diagram for generating a disturbing signal dependent on the speech signal. The jamming of the hearer takes place at predefined time points that are signal-dependent. For this purpose, critical time points (t i,distact ) is determined using three information parameters of the speech signal: the spectral center of gravity "SC" (approximately corresponding to pitch), the short-time energy "RMS" (approximately corresponding to loudness), and the zero-crossing count "ZCR" (to distinguish between speech signal / background noise).

[0060] A set of preselected disturbing signals (e.g. bird calls, chirps, etc.) with associated parameters (SC and RMS) collected by additional preliminary analysis are stored in a digital memory. Suitable disturbing signals preferably have the following properties: on the one hand, they are natural signals familiar to the listener from other situations / everyday life and therefore not related to the signals and contexts to be masked. Furthermore, they are characterized by the fact that they are acoustically characteristic signals of short duration and have as wide a spectrum as possible. Other examples of such signals are the noise of dripping water or the impact of water waves or short wind gusts. Usually disturbing signals are longer than the associated speech segments (e.g. consonants) that completely cover them. It is also possible to store disturbing signals of different lengths and select them to match the duration of the current critical moment.

[0061] A disturbing signal is selected and adapted in time and frequency to the current speech segment. The adapted disturbing signal is then played back to the hearer from a virtual spatial location. For spatialization (BRTF), a short impulse response (256 points) can be used to simulate the outer ear transfer function, so that these disturbing signals are confined to being present as close to the head as possible by the hearer to achieve a strong disturbing effect. For spatial reproduction, multi-channel (e.g., stereo) reproduction is required.

[0062] In the following, the signal processing steps for generating the discrete, spatially distributed, short disturbance signals according to the example embodiment are described in detail. It should be noted that not all steps are always required, and some steps may be performed in different orders, as those skilled in the art will recognize. Also, some calculations may be performed equally well in the frequency domain or the time domain. Some of the processing steps correspond to steps for generating a broadband masking signal, and therefore do not need to be performed a second time in the example embodiment that uses both types of signals for speech obfuscation.

[0063] In section 200, the speech signal sig est is divided into blocks (block length = 512 samples, fs = 44.1 kHz) with a duration of 11.6 ms and a 50% overlap (hop size = 256) (see Section 100).

[0064] These blocks are XBuffer, where n = block index and m = time samples. n From (m), the number of zero crossings per signal block (the zero crossing rate ZCR) is determined in section 205. This can be done using the following formula:

[0065]

number

[0066] In section 210, each signal block is Fourier transformed with NFFT2 = 1024 points (see section 105).

[0067] From these spectra S(k,n) (k=frequency index, n=block index), two further parameters are calculated in sections 215 and 200: the short-time energy (RMS) and the spectral centroid (SC).

[0068]

number

[0069] The courses of the short-time energy RMS and the zero-crossing rate ZCR can also be filtered using signal-dependent thresholds, and regions not meeting these thresholds can be ignored (e.g. set to 0). The thresholds can, for example, be chosen such that a certain percentage of the signal values ​​lie above or below them.

[0070] Each spectrum is spectrally smoothed in section 225 using a recursive first-order discrete-time filter: H(z)=Bs(z) / A(z), where Bs=0.3 and A(z)=1-(Bs-1). * z -1 , in both directions (=acau-sales, second-order zero-phase filter).

[0071] The resulting spectrum is smoothed in time in section 230 using a recursive first-order discrete-time filter: H(z)=Bt(z) / At(z), where Bt=0.3 and At(z)=1-(Bt-1). * z -1 It is.

[0072] For the detection of speech signal sections (onsets) relevant for speech intelligibility (onset detection), first an onset detection function is determined in section 235. For this purpose, the spectrally and temporally averaged spectra are summed across the frequency axis. The resulting signal is logarithmized and time differentiated, with negative values ​​set to zero. To avoid zero values, a regularization (e.g. adding a small number at every frequency grid point) can be performed before logarithmization.

[0073] This onset detection function is scanned for local maxima, which must be at least a specified number of blocks apart. The maxima thus detected can be further filtered using a signal-dependent threshold so that only the most salient maxima remain. The local maxima of the onset detection function thus determined are candidates for perceptually relevant segments of the speech signal to be selectively disrupted using a disturbance signal.

[0074] In the example embodiment, the maximum values ​​of the onset detection functions thus determined in section 240 are checked for validity via a logic unit using the parameters ZCR, RMS and SC. Only if these values ​​are within the defined ranges are these maximum values ​​determined to be the maximum value at the associated critical time t i,distact This means that, for example, at the time of the determined maximum of the onset detection function, the values ​​of RMS, SC and / or ZCR are set to a certain logical condition (e.g. RMS>X1;X2 <SC<X3;ZCR> X4, certain thresholds (X1 to X4) must be satisfied. In the example embodiment, for example, only maxima located in periods that satisfy the above filter conditions for RMS and ZCR (i.e., not in hidden ranges) are considered. The condition that ZCR and RMS must simultaneously satisfy certain threshold conditions can also be used to filter the course of SC by retaining the value of SC when the threshold condition is satisfied and interpolating or extrapolating the inserted value, so that the function SC int is obtained.

[0075] The determined time t i,distact In step 2, one disturbance signal is randomly selected (using section 245) from a selection of N disturbance signals digitally stored in memory 250. Memory 250 contains additional metadata for these disturbance signals: SC and RMS values.

[0076] The selected disturbance signal is divided into blocks in section 255 (with block length 2 and hop size = block length 2 or overlap = 0, respectively, see above) and then Fourier transformed with NFFT2 points in section 260. The parameters of this frequency transform can be different and independent from the version for the speech signal to be masked. Alternatively, the frequency representation of the disturbance signal can be stored directly in the frequency domain.

[0077] The resulting spectrum is then modulated in section 265 using the SC parameter ratios in frequency position (e.g., by single sideband modulation) and / or the RMS parameter ratios in gain to obtain the respective time t i,distact In sig est For this purpose, the onset time t i,distact A ratio is formed between the spectral centroid SC of each speech signal section in and the associated disturbance signal, and the frequency location of the disturbance signal is adjusted to match as closely as possible the frequency location of the speech signal. This is done by subtracting the interpolated spectral centroid SC at the onset time int {t i,distact} function SC int This can be done by comparing the value of with the SC value of the selected disturbance signal and determining a detuning parameter, a positive value of which means an increase in the pitch of the disturbance signal due to single sideband modulation and a negative value results in a decrease in pitch.

[0078] By matching the energy (RMS) of the disturbing signal to the energy of the speech signal portion as well, a predefined ratio of the disturbing signal to the speech signal is achieved. Due to the high effect of reducing speech intelligibility, the disturbing signal can be played at a low volume, so that the increase in the overall sound pressure level at the seating position of the hearing passenger is minimized and no discomfort or disturbance to the passenger is increased or local listening comfort is maintained in the best possible manner.

[0079] In an example embodiment, the resulting corrected spectrum of the disturbed signal is obtained by multiplying the t in section 270 using point-wise multiplication of the frequency grid points (or convolution in the time domain) of the corresponding spectrum. i,distact The signal is spatially variably mapped by a binaural spatial transfer function (BRTF) according to a random direction selection at each time point. Furthermore, in section 275, a direction is randomly selected for the deflection signal. A memory 280 contains the binaural spatial transfer functions (BRTF) corresponding to the possible directions. As described for the masking noise, the spatialization can be performed in the frequency domain or in the time domain. In the time domain, a convolution is performed with the impulse response of the selected outer ear transfer function. The spatialization of the disturbing signal is preferably performed in such a way that the disturbing signal is localized to be present as close to the head as possible by the hearer, thereby achieving a strong scattering effect. Spatialization requires multi-channel (e.g. stereo) reproduction, otherwise a single channel reproduction would suffice, but this is also preferably achieved using two loudspeakers integrated in the headrest.

[0080] In the case of spatialization of the disturbance signal in the frequency domain, the result of the convolution is transformed back into the time domain by an inverse Fourier transform (IFFT) with NFFT2 points in section 285. The inverse transformed time blocks are combined into a time signal using the overlap-add method in section 290. It is clear that the inverse transformation and overlapping of the blocks can be omitted if the preceding steps have already been performed in the time domain.

[0081] The resulting time signal is transmitted to each active neck rest of the hearer. In example embodiments where a masking noise signal is also generated, the masking signal may be summed with the disturbance signal before being output via the speaker of the neck rest.

[0082] The disturbance signal, matched to the speech signal, produces randomly distributed excitation / trigger information that improves the speech target signal without significantly affecting the signal level permanently.

[0083] As mentioned, the signal processing can be performed partly in the frequency domain or in the time domain. The above specific values ​​are only examples of any possible configurations of the frequency transformation and can be modified in many ways. In one possible variant, the energy and frequency matched spectrum (see section 265) is first transformed back to the time domain and then back again to the frequency domain to account for the spatialization, possibly with a different spectral resolution. However, it is also possible to perform the entire processing in the frequency domain. Those skilled in the art of digital signal processing will recognize such variants of the inventive procedure for generating a disturbing signal dependent on a speech signal after considering this disclosure.

[0084] In an example embodiment, both obfuscated signals, the broadband masking noise signal and the disturbing signal, are summed before being output and co-played. The masking noise, which is preferably perceived from the speaker's direction, generates a broadband noise signal adapted to the spectral characteristics of the respective speech segment, on which the short disturbing signals are selectively superimposed (in terms of time and frequency) at particularly relevant points. These disturbing signals, even if played at low volume or low energy, are perceived spatially close to the head and reduce speech intelligibility particularly effectively. However, in combination with the broadband masking noise, the short on and off switching of the disturbing signal is perceived as less disruptive or disturbing. The overall sound pressure level at the seat position of the overhearing passenger is only minimally increased, and the passenger's discomfort or disturbance is not increased or the local listening comfort is maintained as best as possible.

[0085] The above description of the exemplary embodiments has various details that are not essential to the invention defined by the claims. The description of the exemplary embodiments is intended to understand the invention, is purely exemplary, and does not limit the effect on the scope of protection. It will be clear to the skilled person that the described elements and their technical effects can be combined with each other in different ways, resulting in further exemplary embodiments covered by the claims. Furthermore, the described technical features can be used in devices and methods implemented, for example, by programmable devices. In particular, they can be implemented by hardware elements or by software. As is known, the implementation of digital signal processing is preferably carried out by specially designed signal processors. The communication between the individual parts of the described devices can occur by wire (for example, a bus system) or wirelessly (for example, Bluetooth or Wifi). Protection is also expressly claimed for computer-implemented implementations and for associated programs or machine codes in the form of a data carrier or in a downloadable representation.

Claims

1. 1. A method of masking a speech signal in a zone-based audio system, comprising: detecting a speech signal to be masked in a voice zone; converting the detected speech signal into spectral bands; rectifying the spectral values ​​of at least two spectral bands; generating a noise signal based on the rectified spectral values; and outputting the noise signal as a masking signal for the speech signal in another audio zone.

2. generating a noise signal based on the rectified spectral values; generating a broadband noise signal; transforming the generated noise signal into the frequency domain; multiplying the frequency representation of the noise signal by a frequency representation of the speech signal while taking into account the rectified spectral values; The method of claim 1 , comprising:

3. The method of claim 2 , wherein the frequency representation of the speech signal is generated by interpolating the spectral values ​​of the bands following rectification of the spectral values.

4. estimating a background noise spectrum; comparing the spectral values ​​of the speech signal with the background noise spectrum; considering exclusively spectral values ​​of the speech signal that are greater than corresponding spectral values ​​of the background noise spectrum; The method of claim 1 further comprising:

5. 2. The method of claim 1, wherein the detected speech signal is transformed into spectral bands of blocks of the speech signal and by a Mel filter bank, and optionally a temporal smoothing of the spectral values ​​of the Mel bands is performed.

6. 2. The method of claim 1, wherein the noise signal is spatially represented at the output by multi-channel reproduction, preferably by multiplication by the binaural spectrum of an acoustic transfer function.

7. 7. The method of claim 6, wherein the noise signal is spatially output in the other audio zones so as to appear to emanate from the dominant direction of the speaker of the speech signal to be masked.

8. determining a time point in the speech signal that is related to speech intelligibility; generating a disturbance signal at the determined time point; outputting the disturbance signal at the determined time as a separate masking signal for the other audio zone; The method of claim 1 further comprising:

9. 9. The method of claim 8, wherein the time points associated with speech intelligibility are determined using extrema of a spectral function of the speech signal, the spectral function being determined based on a summation of spectral values, optionally averaged, over the frequency axis.

10. 9. The method of claim 8, wherein the time points related to speech intelligibility are verified using parameters of the speech signal such as zero-crossing rate, short-time energy and / or spectral centroid.

11. 9. The method of claim 8, wherein the disturbing signal at the particular time point is randomly selected from a set of predetermined disturbing signals and / or is matched to the speech signal in terms of its spectral characteristics and / or its energy.

12. 1. A method of masking a speech signal in a zone-based audio system, comprising: detecting a speech signal to be masked in a voice zone; determining a time point in the speech signal that is related to speech intelligibility; generating a disturbance signal for said determined time point, said disturbance signal being adapted to said speech signal in terms of spectral characteristics and / or its energy; and outputting said disturbing signal as a masking signal in another audio zone at said particular time.

13. 13. The method of claim 12, wherein the time points associated with speech intelligibility are determined using extrema of a spectral function of the speech signal, the spectral function being determined based on a summation, optionally averaged, of spectral values ​​on the frequency axis.

14. The method of claim 12 , wherein the time points associated with speech intelligibility are verified using parameters of the speech signal such as zero-crossing rate, short-time energy and / or spectral centroid.

15. The method of claim 12 , wherein the disturbance signal at the particular time point is randomly selected from a set of predetermined disturbance signals.

16. converting the captured speech signal into spectral bands; rectifying the spectral values ​​of at least two spectral bands; generating a noise signal based on the rectified spectral values; outputting the noise signal as an additional masking signal for the speech signal in the other audio zone; The method of claim 12 further comprising:

17. generating a noise signal based on the rectified spectral values; generating a broadband noise signal; transforming the generated noise signal into the frequency domain; multiplying the frequency representation of the noise signal by a frequency representation of the speech signal while taking into account the rectified spectral values; 17. The method of claim 16, comprising:

18. estimating a background noise spectrum; comparing the spectral values ​​of the speech signal with the background noise spectrum; considering only spectral values ​​of the speech signal that are greater than corresponding spectral values ​​of the background noise spectrum; 17. The method of claim 16, further comprising:

19. 17. The method of claim 16, wherein the conversion of the captured speech signal into spectral bands is for blocks of the speech signal and is performed using a Mel filter bank, optionally with temporal smoothing of the spectral values ​​for the Mel bands.

20. 20. The method according to any one of claims 1 to 19, wherein the masking signal is spatially represented at the output using multi-channel reproduction in the other sound zones, preferably by multiplication by the binaural spectrum of an acoustic transfer function.

21. 21. The method of claim 20, wherein the masking signal is spatially output in the other sound zones so as to appear to emanate from random directions and / or from near the listener's head in the other sound zones.

22. 1. An apparatus for generating a masking signal in a zone-based audio system, the apparatus receiving a speech signal to be masked and generating the masking signal based on the speech signal, comprising: means for converting the detected speech signal into a spectral band; means for rectifying the spectral values ​​of at least two spectral bands; means for generating a noise signal as a masking signal based on the rectified spectral values.

23. means for determining points in the speech signal that are related to speech intelligibility; means for generating a disturbance signal for said time point; means for adding the noise signal and the disturbance signal and outputting the resulting signal as a masking signal; 23. The apparatus of claim 22, further comprising:

24. 1. An apparatus for generating a masking signal in a zone-based audio system, the apparatus receiving a speech signal to be masked in an audio zone and generating a masking signal based on the speech signal, the apparatus comprising: means for determining points in the speech signal that are related to speech intelligibility; means for generating a disturbance signal for said relevant time instant, said disturbance signal being adapted to said speech signal in terms of its spectral characteristics and / or its energy; means for outputting the disturbing signal as a masking signal at the specific time in another audio zone; An apparatus comprising:

25. means for converting the detected speech signal into a spectral band; means for rectifying the spectral values ​​of at least two spectral bands; means for generating a noise signal as a masking signal based on the rectified spectral values; means for adding the noise signal and the disturbance signal and outputting the resulting signal as a masking signal; 25. The apparatus of claim 24, further comprising:

26. Apparatus according to any one of claims 22 to 25, further comprising means for generating a multi-channel representation of the masking signal to allow spatial reproduction of the masking signal.

27. 1. A zone-based audio system comprising a plurality of audio zones, 25. A zone-based sound system, wherein one sound zone comprises at least one microphone for detecting speech signals and another sound zone comprises at least one loudspeaker, the microphones and loudspeakers preferably being located in headrests of passenger seats in a vehicle, the sound system comprising an apparatus for generating a masking signal as claimed in claim 22 or 24, the apparatus receiving a speech signal from the microphone of the one sound zone and transmitting the masking signal to the one or more loudspeakers of the other sound zone.