Device and method for adaptive, harmonic voice masking sound generation
Patent Information
- Application Number
- EP2023798818
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-07
- Filing Date
- 2023-11-03
- Publication Date
- 2025-09-17
AI Technical Summary
Existing noise masking systems in open-plan offices are inflexible and unpleasant for users, failing to adapt to changing background noise conditions and prioritize acoustic privacy and listening horizon needs, leading to cognitive performance decline and user dissatisfaction.
A device and method for adaptive, harmonic speech masking sound generation that analyzes frequency band-limited signal components to dynamically adjust masking sound levels and frequency spectrum, using a control algorithm to ensure the masking sound is only as loud as necessary, with a harmonic component for improved user acceptance.
The solution effectively improves cognitive performance and user satisfaction by providing a pleasant and secure masking sound that adapts to changing noise conditions, minimizing cognitive decline and maximizing acoustic satisfaction.
Smart Images

Figure 1.1
Abstract
Description
[0001] Apparatus and method for adaptive, harmonic speech masking sound generation
[0002] Description
[0003] The application relates to noise masking, in particular speech masking, and, more specifically, to an apparatus and a method for adaptive, harmonic speech masking sound generation.
[0004] In particular, speech sounds and unpredictable noises with peak levels that contrast significantly with the background level impair our cognitive performance (see Bodin Danielsson & Bodin, 2009). These negative effects on visual and auditory short-term memory of noise, such as speech, are called the irrelevant sound effect (ISE).
[0005] Requirements for an acoustic workplace environment vary both over the course of the workday and depending on the tasks employees perform. For example, people working in a crowded office have a strong need for acoustic privacy, while people working in a sparsely occupied office may need an expanded auditory horizon, for example, to avoid being surprised by the sudden appearance of other people (Zuydervliet et al., 2008).
[0006] To counteract the interference caused by ISE, some open-plan offices employ the approach of global masking systems via loudspeakers. This involves playing broadband static noise over central loudspeakers throughout the office spaces with the goal of masking background noise. The resulting reduction in the signal-to-noise ratio of, for example, disturbing speech signals to the now increased background noise minimizes speech intelligibility (Zuydervliet et al., 2008). However, such noise can be perceived as unpleasant by users and therefore rejected (see Keus Van De Poll, Marijke et al., 2015).
[0007] Some approaches are based on dynamically adjusting the volume of masking sounds to changing background noise conditions. For example, there are global systems in which the masking sound changes at set time intervals or based on microphone measurements across the entire office. These offer a rather inflexible and therefore inadequate solution. Furthermore, an increase in employee satisfaction has been observed when employees have the opportunity to personalize their workplace (see Huang, Robertson, & Chang, 2004; Lee & Brand, 2010).
[0008] For a sound to have sufficient masking effect, it must have a sufficiently broadband noise signal in all frequency ranges where the interfering signal occurs. Pink noise, for example, which has a volume drop of approximately 3 dB per octave, has been identified as an effective masking signal. However, it is often subjectively perceived as disturbing by subjects and therefore tends to be rejected (Schlittmeier & Hellbrück, 2009). Various other studies have proposed and investigated masking signals with other spectra. Signals such as frequency-matched pink noise, speech-like murmurs from multiple speakers, or even natural signals such as the sound of spring water were considered (Hongistob et al., 2017; Veitch et al., 2002; Wang, Drotleff & Li, 2012). Natural sound sources appear to significantly improve user acceptance (Haapakangas et al., 2011).Music has also been investigated as a source of task-irrelevant sounds in studies, but this turned out to be less effective (Haapakangas et al., 2011; Schlittmeier & Hellbrück, 2009).
[0009] To incorporate the subjective judgments described above in addition to the sole psychoacoustic effectiveness, Chanaud (2007) presented two systems of adaptive sound masking. First, a time-based system in which the sound pressure level of the masking sound varies at static time intervals throughout the day. This requires predicting different acoustic privacy needs and the likely level of noise disturbance at different times of day. For example, it is important for employees to be able to hear the presence of other people overnight and in the early morning. Accordingly, no masking sound, or only a very quiet one, would suffice at these times. During peak hours, when most employees arrive at or leave their workplace, there is a lot of social interaction, which can lead to a high degree of acoustic distraction, which should be compensated for by increasing the volume of the masking sound.During the lunch break, the need for masking is again low. The 10th percentile (L10) of the measured sound level describes the sound level that was reached at least 10% of the time period under consideration. The 90th percentile (L90) describes the sound level that was reached at least 90% of the time period under consideration. Zuydervliet et al. (2008) also propose that the 10th and 90th percentile values should be determined for adaptive masking sound control. The 90th percentile represents L. A F,9O% representative of the background noise of the ambient sound and the 10th percentile L A F, % describes the activity transients of noise in the background noise condition. The difference between these L AF,10%-90% values thus describe an SNR of interference components and background noise (Zuydervliet et al., 2008). If the SNR is high, the background noise has a strong changing-state character and therefore causes an ISE. The determined L AF According to Zuydervliet et al. (2008), the ,io%-9o% value should be compared with a target value L AF ,io%-9o%,ziei, i.e. with an optimal percentile value difference. If the difference of the determined L AF,IO%- 90% greater than the target value, the masking sound level should slowly increase and thus reduce the signal-to-noise ratio (SNR) of the total sound. If the difference is smaller, the SNR of the sum signal is already smaller than the minimum required to prevent ISE, and the masking sound can slowly become quieter. Additional parameters such as a weighting factor W, a maximum volume change per minute, and a parameter for adjusting the sensitivity can thus influence the optimal volume of the masking sound (Zuydervliet et al., 2008). The target value (L AFThe sensitivity of the percentiles (i.e., io%-9o%,ziei) should be between 3 and 10 dB, while the weighting factor should be between 0.5 and 4. The time period over which the analysis of the percentile values takes place determines the sensitivity of the system; a period of 15 seconds is suggested for this purpose. If a longer period is selected, level fluctuations are less significant, and the control system reacts more slowly to changing sound conditions. A maximum rate of change of 0.05 dB per second is specified (L'Esperance et al., 2017).
[0010] According to Chanaud (2007), the level increase should generally occur faster than the level decrease of the masking sound. Furthermore, it is suggested that the upper and lower limits of adaptive masking sound systems be limited. This should ensure that sufficient masking is ensured at all times, while at the same time, a maximum acceptable level is not exceeded. Zuydervliet et al. (2008) suggests a dynamic range of 5 dB for the masking sound volume, while L'Esperance (2017) proposes 3 dB.
[0011] While the majority of previous approaches focus on global masking systems, in which the masking sounds are played over central loudspeakers in, for example, an open-plan office, Schlittmeier and Hellbrück (2009) recommend local masking systems for individual workers. However, since neighboring employees can be disturbed by crosstalk from individual masking sounds played over loudspeakers at each workstation, playing these masking sounds over headphones in combination with adaptive level control is a promising approach. The acceptance of such a headphone-based masking system should also be improved by the personalization options compared to loudspeaker masking systems. Various studies have shown that satisfaction with the work environment is increased when it can be controlled by employees (Huang et al., 2004; Lee & Brand, 2010).
[0012] However, not only employee satisfaction, but also their physical health and performance should be increased through more control over the work environment (see Cohen, 1980; Quick, 1990). Renz (2019) also focuses on percentile value differences L A F,10% - LAF,90% and has developed a new method for predicting an expected decrease in performance (DP). Renz (2019) suggests 2 to 3 dB as a suitable target value (L AF ,io%-9o%,ziei) for an adaptive level control of masking sounds.
[0013] The DP values evaluated and plotted by Renz (2019), “Personalised sound masking in open offices. A trade-off between annoyance and restoration of working memory performance?” Stuttgart: Fraunhofer Verl, Stuttgart, as a function of the prediction parameter L A,IO%-9O% are shown in Renz, 2019 on page 204. Renz, 2019, shows on page 204 a plot of the cognitive performance prediction model of the resulting DP with the prediction parameter LAF 10-90.
[0014] US 2003 / 103632 A1 shows an adaptive noise masking system and method that divides unwanted noise into time blocks and estimates the frequency spectrum and power level, while continuously generating white noise with an appropriate spectrum and power level to mask the unwanted noise.
[0015] CN 110362789 A shows a noise masking method and an adaptive noise masking system with a noise masking database, a noise satisfaction agent model and a self-adapting noise masking search system.
[0016] US 2015 / 194144 A1 shows a multi-microphone subsystem to detect sounds, a spectrum analyzer to determine a power characteristic of the detected sound, and a spatial analyzer to detect a directional characteristic of the sound.
[0017] An apparatus according to claim 1, a method according to claim 14 and a computer program according to claim 15 are provided.
[0018] A device for generating speech masking sound according to one embodiment is provided. The device comprises an analyzer for analyzing each frequency-band-limited signal component of a plurality of frequency-band-limited signal components of a microphone signal during an analyzed period of time in order to obtain information about the frequency-band-limited signal component. Furthermore, the device comprises a masking signal generator for generating a masking signal depending on the information about the frequency-band-limited signal component of each of the plurality of frequency-band-limited signal components. The information about the frequency-band-limited signal component depends on a first sound level that was reached at least during a first time period during the analyzed period of time.Furthermore, the information about the frequency-band-limited signal component depends on a second sound level reached at least during a second period of time during the analyzed period, the second period of time being different from the first period of time.
[0019] Furthermore, a method for generating speech masking sound according to one embodiment is provided. The method comprises:
[0020] Analyzing each frequency-band-limited signal component of a plurality of frequency-band-limited signal components of a microphone signal during an analyzed period to obtain information about the frequency-band-limited signal component. And:
[0021] Generating a masking signal depending on the information about the frequency band-limited signal component of each of the plurality of frequency band-limited signal components.
[0022] Furthermore, a computer program with a program code for carrying out the method described above is provided according to one embodiment.
[0023] The information about the frequency-band-limited signal component depends on a first sound level that was reached at least during a first time period during the analyzed period. Furthermore, the information about the frequency-band-limited signal component depends on a second sound level that was reached at least during a second time period during the analyzed period, wherein the second time period is different from the first time period.
[0024] Embodiments provide a control algorithm that can adjust a speech masking signal for presentation via headphones in a way that is both pleasant and secure and psychoacoustically validated.
[0025] The DP values described above, evaluated and plotted by Renz (2019) as a function of the prediction parameter LAF, %-9O%, which can be read in Renz, 2019 on page 204, form a scientific basis of considerations on which embodiments are based.
[0026] Some versions provide masking sound, which can be individually adjusted via headphones, for example, and is only used to the extent needed. This creates an effective way to improve both cognitive performance and employee satisfaction in the workplace.
[0027] Preferred embodiments of the invention are described below with reference to the drawings.
[0028] The drawings show:
[0029] Fig. 1 shows a device for speech masking sound generation according to a
[0030] Embodiment.
[0031] Fig. 2 shows a signal flow diagram with a frequency division into nine
[0032] Octave bands according to one embodiment.
[0033] Fig. 3 shows a signal flow diagram for adaptive noise masking according to one embodiment.
[0034] Fig. 4 shows a signal flow diagram of a control value checker according to one embodiment. Fig. 5 shows a control loop according to one embodiment.
[0035] Fig. 1 shows a device for generating speech masking sound according to an embodiment.
[0036] The device comprises an analyzer 110 for analyzing each frequency-band-limited signal component of a plurality of frequency-band-limited signal components of a microphone signal during an analyzed period of time in order to obtain information about the frequency-band-limited signal component.
[0037] Furthermore, the device comprises a masking signal generator 120 for generating a masking signal depending on the information about the frequency-band-limited signal component of each of the plurality of frequency-band-limited signal components.
[0038] The information about the frequency-band-limited signal component depends on a first sound level that was reached at least during a first time period during the analyzed period. Furthermore, the information about the frequency-band-limited signal component depends on a second sound level that was reached at least during a second time period during the analyzed period, wherein the second time period is different from the first time period.
[0039] According to one embodiment, analyzer 110 can be configured, for example, to determine a microphone signal sound level difference between the first sound level and the second sound level for each signal component of the plurality of frequency-band-limited signal components of the microphone signal. In this case, masking signal generator 120 can be configured, for example, to determine the masking signal depending on the microphone signal sound level difference of each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the microphone signal.
[0040] In one embodiment, the masking signal generator 120 can, for example, be configured to determine the masking signal by determining, for each signal component of the plurality of frequency-band-limited signal components, a level value for a frequency-band-limited component of the masking signal corresponding to a frequency range of this signal component, depending on the microphone signal-sound level difference of this signal component, and to perform a level adjustment of this frequency-band-limited component of the masking signal using this level value. According to one embodiment, the analyzer 110 can, for example, be configured to determine an overall signal that depends on the microphone signal but is different from the microphone signal. In this case, the analyzer 110 can, for example,be configured to determine an error value for each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the overall signal, which error value indicates a difference between a target value for a total signal sound level difference and a current total signal sound level difference of the overall signal. The masking signal generator 120 can, for example, be configured to determine the masking signal depending on the error value for each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the overall signal. The analyzer 110 can, for example,be designed to determine the current total signal sound level difference between a third sound level and a fourth sound level for each frequency-band-limited signal portion of the plurality of frequency-band-limited signal portions of the total signal, wherein the third sound level is a sound level that was reached at least during a third time period during an analyzed time period in the frequency-band-limited signal portion of the total signal, and wherein the fourth sound level is a sound level that was reached at least during a fourth time period during the analyzed time period in the frequency-band-limited signal portion of the total signal.
[0041] In one embodiment, the analyzer 110 may, for example, be configured to determine each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the overall signal depending on a feedback time segment of the masking signal to this frequency-band-limited signal component.
[0042] According to one embodiment, the analyzer 110 can, for example, be configured to determine each of the plurality of frequency-band-limited signal components of the overall signal depending on an attenuation factor for this frequency-band-limited signal component of the overall signal, wherein the analyzer 110 is configured to apply the attenuation factor for this frequency-band-limited signal component to the corresponding frequency-band-limited signal component of the microphone in order to obtain an attenuated microphone signal for this frequency-band-limited signal component.
[0043] In one embodiment, the analyzer 110 may, for example, be configured to determine each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the overall signal as the sum of the feedback time portion of the masking signal to this frequency-band-limited signal component and the attenuated microphone signal to this frequency-band-limited signal component.
[0044] According to one embodiment, the masking signal generator 120 can, for example, be designed to determine the masking signal as a function of a correction value for each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the overall signal, wherein the masking signal generator 120 is designed to determine the correction value for this frequency-band-limited signal component as a function of the error value for this frequency-band-limited signal component.
[0045] In one embodiment, the masking signal generator 120 may further be configured, for example, to determine the correction value for this frequency-band-limited signal component as a function of a temporal predecessor value of this correction value.
[0046] According to one embodiment, the masking signal generator 120 can be designed, for example, to determine the correction value as a function of z n = z n- + s^n(e) * ge), where z n is the correction value at a time n, where z n- is the temporal predecessor value of this correction value at a time n-1, where e is the error value, where sgn denotes a sign function, and where g(e) denotes a function dependent on the error value e.
[0047] In one embodiment, the masking signal generator 120 can, for example, be designed to determine the masking signal as a function of a control value for each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the microphone signal, wherein the masking signal generator 120 is designed to determine the control value for this frequency-band-limited signal component as a function of the microphone signal sound level difference of this frequency-band-limited signal component and as a function of the error value and the correction value of this frequency-band-limited signal component of the overall signal.
[0048] According to one embodiment, the masking signal generator 120 can, for example, be designed to determine the control value for this frequency-band-limited signal component by forming a sum of the microphone signal sound level difference of this frequency-band-limited signal component and the error value and the correction value of this frequency-band-limited signal component of the overall signal.
[0049] In one embodiment, the masking signal generator 120 can, for example, be configured to determine the level value for a frequency-band-limited component of the masking signal as a function of the control value for this frequency-band-limited signal component and as a function of a previous level value for this frequency-band-limited component of the masking signal.
[0050] Embodiments provide a control algorithm that enables a masking sound to be dynamically adapted in volume and frequency spectrum to a background sound condition.
[0051] The algorithm can, for example, independently detect the extent to which background noise may have a disruptive influence on cognitive performance. A microphone signal is used to assess the noise level.
[0052] The algorithm works on various devices using the available technology. Since it cannot be assumed that all devices are equipped with calibrated, standard-compliant microphones that meet the requirements for sound level meters according to DIN EN 61672-1, the algorithm does not require any knowledge of the absolute sound pressure level. The algorithm determines the 90% and 10% percentile values as control parameters.
[0053] A masking sound is generated and played continuously, providing sufficient masking power to prevent a potential cognitive decline caused by the background noise condition. The algorithm has a suitable sensitivity to the background noise condition to ensure that spontaneously occurring noises that are not representative of the background noise condition are not used for control.
[0054] The masking sound generated by the algorithm is only as loud as necessary at any given time. The goal is not only to increase performance, but also to ensure the acoustic satisfaction of the users. Masking sounds are generally perceived as more unpleasant than silence. To ensure the greatest possible user acceptance, the algorithm constantly detects the minimum masking sound level required and continuously uses this as the target value for the level control. The target value is the ratio of the L A F, % percentile value to L AThe F90% percentile value is used. The masking sound adapts to the frequency spectrum of the background noise. The control times with which the masking sound is controlled in volume are selected by the algorithm so that volume fluctuations are barely noticeable. This is useful so that the masking sound itself does not distract the user. At the same time, volume changes occur quickly enough to be able to react to changing acoustic conditions in the background noise.
[0055] To further improve acceptance, the algorithm adds a harmonic component to the masking sound, which ensures a pleasant sound of the masking sound.
[0056] The algorithm according to one embodiment is described in detail below.
[0057] Thus, Fig. 2 shows a signal flow diagram according to an embodiment with a frequency division into nine octave bands, which in the example of Fig. 2 have center frequencies at 63 Hz, 125 Hz, 250 Hz, 500 Hz, 1000 Hz, 2000 Hz, 4000 Hz, 800 Hz, 1600 Hz.
[0058] Fig. 2 illustrates the part of the algorithm in which the frequency division of the microphone input signal into bands, for example, octave bands, takes place. Fig. 2 also shows how the various band-filtered masking sound components are mixed together and adjusted with a calibration factor W before the masking signal is output to the headphones. The light blue-framed elements labeled "Adaptive Level Control" in Fig. 2 represent the part of the algorithm illustrated in detail in Fig. 3.
[0059] Fig. 3 shows a signal flow diagram for adaptive noise masking according to one embodiment.
[0060] This part of the algorithm performs the level measurement, percentile difference determination, and continuous calculation of the control value u. Furthermore, the control values u are smoothed using set control times to obtain the level value p, which in turn controls the volume of the respective band-filtered masking sound component. The part of the algorithm in which the current LA,F, %-90%,Ges value is calculated is also outlined accordingly (blue) in Fig. 3 and is clearly illustrated in Fig. 4.
[0061] Thus, Fig. 4 shows a signal flow diagram of a control value checker according to one embodiment. In particular, Fig. 4 illustrates how the percentile values of the level values of the entire signal (microphone signal * damping factor + (returned) masking noise component) are calculated. This resulting calculated level difference (LAF,10% - 90%,Total) is compared with the target value to obtain an error value (e).
[0062] Fig. 5 shows a control loop according to one embodiment. In particular, Fig. 5 shows how a correction value is calculated from the error value e.
[0063] The algorithm's input, as shown in Fig. 2, is the digital audio signal from a microphone continuously recording ambient sound. This signal is first divided into bands by octave filters according to DIN EN 61260-1, e.g., nine octave bands (e.g., with center frequencies in the range 63 Hz - 16,000 Hz). The band-limited signals are analyzed and processed in the respective signal paths. Since adaptive level control is to be performed per band, the masking signal is also divided into individual bands. The volume of these band-filtered masking sounds is individually controlled in the signal paths (corresponding to the respective octave band), and then mixed together again to form an overall masking sound.This allows the current interference influence to be calculated for each octave band, which is used to determine the volume control that the respective frequency range of the masking sound should have in order to sufficiently mask the interference sounds occurring across the entire frequency spectrum.
[0064] The controlled masking signal is supplemented with an additional harmonic component. The harmonic component is a type of music that improves the acceptance and subjective perception of the masking sound. The harmonic component is included in the calculation of the expected interference effect of the acoustic environment described below. Furthermore, the harmonic component, mixed with the controlled masking component, is played through the headphones. The harmonic component is psychoacoustically validated, i.e., its suitability has been tested in listening tests (no changing-state behavior, as the ISE is not triggered). The harmonic component can, for example, be an uncompressed stereo file (StereoFile), which is played and controlled by the algorithm. The input for adaptive sound masking (see Fig. 3) is the band-filtered, A-weighted audio signal from the microphone.For audible frequencies, A-weighting is a commonly used frequency weighting that represents the ear's response to sound pressure or loudness. Regarding temporal weighting, the weightings F (fast), S (slow), and I (impulse) indicate how quickly a response occurs to a change in sound level. LAF denotes a sound level with A-frequency weighting and F-time weighting. The percentile levels L. A F, % and LAF,90% indicate which levels were reached in 10 % and 90 % of the measuring time, respectively.
[0065] The equivalent continuous sound level of the said audio signal is now determined in accordance with DIN EN 61672-1. For this purpose, a root mean square (RMS) is first determined for each sample. The level values are then integrated over 125 milliseconds. To avoid possible errors in further signal processing due to unrealistically small amplitude values, the values are limited by a minimum amplitude value. The measured level values are stored in a continuous list, with the list length defining the observation period over which the percentile values are analyzed. Due to the previous level measurement, the list is supplemented with a new level value every 125 milliseconds, and an old value is deleted. A percentile value calculation (L A F,90%, and L A F, %) in the list. Then the value of the 90th percentile L AF ,9o%, of that of the 10th percentile L AF, % subtracted to determine the percentile value difference of the background condition L AF ,10%-90%, HSB ZU (HSB = background noise condition). Since the list of LAF values is updated every 125 milliseconds, a new L is also calculated every 125 milliseconds. AF ,IO%-9O%, HSB value.
[0066] From the difference between these continuously determined percentile values, level differences are calculated, which can be associated with the decrease in performance, as the previously described study by Renz et al. (2018) shows. The higher the relative level of the activity transients L AF .IO%, the greater the distraction (Zuydervliet et al., 2008). To counteract this ISE-induced performance degradation, the background noise level L AF ,9o%, can be increased by adding a masking noise to such an extent that the level difference to the activity sound level L AF,10% is sufficiently reduced. The control described below ensures that the difference between these two values is as small as possible (e.g., less than 3 or, for example, between 2 and 3. Other target values can also be selected). To implement a control that ensures that a LAF, %-90% target value is achieved, the total signal from masking sound and background sound is examined for its LAF, %-90% value.
[0067] Fig. 3 illustrates the part of the algorithm whose functionality is described below. The masking sound is played through headphones, which means that an analysis of the actual percentile values at the position of the user's ear should be performed. However, an exact analysis of this sound condition would only be possible using a microphone located in the ear cup. ANC headphones usually have such a microphone built into the cups, but the microphone signal cannot be used without knowledge of the headphones' integrated signal processing, if it is even possible to capture this signal at all.
[0068] In some embodiments, the algorithm is intended to be universally usable with headphones without ANC. Therefore, in such embodiments, the signal arriving at the user's ear is estimated. If access to the microphone is possible, the value can also be determined directly. Subsequent control is then performed using the measured value instead of the estimated one, but is otherwise identical.
[0069] For the estimation, the extent to which the background noise is reduced by the headphones used should be known. The determined equivalent continuous sound levels of the background noise condition are offset against the attenuation factor in this part of the algorithm to obtain an estimated relative sound level of the background noise condition at the position of the user's ear. A level measurement is now also performed on the masking sound signal (and the harmonic component), which was tapped after its level adjustment (see Fig. 4). The values thus determined are added to the background noise multiplied by the attenuation factor, resulting in the estimated relative total sound level L AF ,IO%-9O%, total can be determined.
[0070] For example, a person using an implementation of the adaptive masking signal generator algorithm in hardware or software can adjust the masking sound generator's playback volume for the current sound condition at the beginning of use. This is done, for example, via a fader in the graphical user interface or a potentiometer on the headphones, which controls the calibration factor W. The calibration factor is applied independently of the level control at the end of the signal path. To calculate the self-correcting manipulated variable u, the algorithm analyzes the microphone input signal and calculates level differences of LAF, % and LAF,90% (see Fig. 3). These level differences LAF, %-90%,HSB are intended to control the masking sound's volume per octave band.
[0071] However, the relationship between the prediction parameter LAF, %-90% and a prospective DP value is not linear (Renz, 2019). This means that simply increasing the volume of the masking sound by the determined LAF, %-90%,HSB value does not necessarily adequately mask the disturbing sound components of the HSB. To verify whether the signal reaching the user's ear (attenuated ambient sound + masking signal) can truly be considered a disturbance-free sound condition, it is analyzed as described above to determine the total value of the percentile differences L A F,io%-9o%,Ges to be determined.
[0072] In the next step, LAF,io%-9o%,Ges is to be compared with a target value LAF, %-9o%,ziei. A suitable target value at which a performance degradation just does not occur significantly is between 2 dB and 3 dB. For example, a target value LAF,io%-9o%,ziei of 2.5 dB is used for the algorithm. This leads to a target value range between 2 dB and 3 dB, within which L A F,io%-9o%,Ges is moved. However, the target value can also be chosen differently. When comparing the percentile value differences, the error value e describes the difference between LAF, %-9o%,ziei and LAF, %-9o%,Ges. The manipulated variable u, which controls the volume of the masking sound, is defined as the sum of LAF, %-90%,HSB and a correction value z (see equation 1). u = LAF,io%-90%,HSB + z (1)
[0073] For constant LAF, %-90%,HSB values and a positive error value, the correction value z must increase until an error value of 0 is reached. As soon as the error value drops below 0, the correction value must continuously decrease again. The correction value will increase and decrease again until it reaches a value at which the error value remains constant at 0. However, the correction value z should increase or decrease more slowly the closer the error value gets to 0. Since there is a tolerance range of + / - 0.5 dB around LAF,io%-9o%,zii, z can constantly increase or decrease until the tolerance limit is reached. From an error value of 0.5, z should change in smaller steps the closer e approaches 0. This is to prevent the error value from being corrected beyond the zero point due to excessive correction.In extreme cases, this could lead to the control system endlessly oscillating the correction value between two extremes in the positive and negative value range. For this reason, the current correction value zn is calculated from the last correction value zn-1 , plus or minus a correction factor g{e). This correction factor depends on the size of the error value e and is clearly defined for different conditions (see equation 3). This part of the algorithm is shown in Fig. 5. n = z n -i + sgn(e) * g(e) (2) 0.5 if |e| > 0.5 0.1 if 0.5 > |e| > 0.25
[0074] (3) 0.05 if 0.25 > |e| > 0.1 0.01 if 0.1 > |e|
[0075] Masking sounds should generally have a maximum sound pressure level between 45 dß(A) and 48 dB(A). This maximum is justified by the fact that higher sound levels over a longer period of time are usually perceived as extremely disturbing (Haapakangas et al., 2011). Therefore, the algorithm limits the upper and lower limits of u. However, the masking system described in this invention report cannot record absolute sound levels, which is why the maximum possible sound level values are controlled by the user's own calibration. The dynamic range of the adaptive masking signal is set to 26 dB, but can be changed depending on the implementation. This means that even with a slightly disturbing HSB, the L A F,io%-9o%,ziei value can be achieved, with the masking sound being as quiet as possible.
[0076] A key requirement for the algorithm is that the HSB is sufficiently masked at all times. However, a change in volume must not cause the masking sound itself to become a disturbing factor. This is because a changing-state character occurs, among other things, when there is strong variability in the amplitude response (Liebl, 2006). And sounds with a changing-state character have a negative effect on cognitive performance. For this reason, the manipulated variable u does not directly regulate the level of the masking sound. By inserting a time ramp, a smoothing of the level values can be achieved. A time ramp ensures a continuous approximation between an old manipulated variable u and a new one. n -i and a new u n . The attack time t Ata£ k, in which the level should increase from an old to a new value, as well as the release time t ReThe rate at which the level drops again can be set separately. Equation 4 describes the current level value p n , which is determined by the current input value u n (Setpoint) the last output level value p n -i and t At ack and t Re iease, defined. The time parameters t At ack, or t Re iease, result from the input sample rate and the desired attack and release times.
[0077] A balance must be struck between whether it's more important for the masking sound to reach a sufficient volume as quickly as possible, or whether a volume change that's as unnoticeable as possible is a priority. Five seconds is suggested as a compromise.
[0078] However, the moment the HSB level falls within the observation period, a high L AF,10%-90%,HSB value, which in turn would lead to a strong increase in the level of the masking signal. This level increase, due to its changing-state character, can in turn represent a distraction without masking any noise. To prevent a level increase in such a situation, the determined L A F,90% values are continuously monitored for strong fluctuations. If a newly arriving L AF ,9o% value compared to the last L AF ,9o% value by more than 2 dB, the attack time t Ata£ k of the time ramp is set to 90 seconds for the duration of a viewing period (5 seconds). An attack time of this magnitude means that no noticeable level increase is possible. After the five seconds have elapsed, the attack time is reset to its regular value, and the level can be adjusted normally again.
[0079] After level adjustment, the entire adaptive harmonic speech masking signal (consisting of the masking and harmonic components) is reproduced via the audio output of the terminal device (digital or analog) via headphones.
[0080] In embodiments, a masking signal is provided that is both pleasant and effective, and reliably achieves psychoacoustically determined target values within a specified time interval, the correlation of which with cognitive performance, for example, is known. Embodiments are based on the control system adjusting the masker by means of an estimation depending on the expected interference effect.
[0081] For example, certain designs can be used in office spaces, especially in offices for multiple people, and can be specifically adapted for use with headphones. Other areas of application could include medical use, therapy, or even tourism.
[0082] Although some aspects have been described in the context of a device, it should be understood that these aspects also represent a description of the corresponding method, so that a block or component of a device can also be understood as a corresponding method step or as a feature of a method step. Analogously, aspects described in the context of or as a method step also represent a description of a corresponding block, detail, or feature of a corresponding device. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, some or more of the key method steps may be performed by such an apparatus.
[0083] Depending on specific implementation requirements, embodiments of the invention may be implemented in hardware or in software, or at least partially in hardware or at least partially in software. The implementation may be carried out using a digital storage medium, for example a floppy disk, a DVD, a Blu-ray disc, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a FLASH memory, a hard disk, or other magnetic or optical storage device on which electronically readable control signals are stored that can interact or interact with a programmable computer system such that the respective method is carried out. Therefore, the digital storage medium may be computer-readable.
[0084] Some embodiments according to the invention thus comprise a data carrier having electronically readable control signals capable of interacting with a programmable computer system such that one of the methods described herein is carried out. In general, embodiments of the present invention can be implemented as a computer program product with program code, wherein the program code is effective to carry out one of the methods when the computer program product is executed on a computer.
[0085] The program code can, for example, also be stored on a machine-readable medium.
[0086] Other embodiments include the computer program for performing one of the methods described herein, wherein the computer program is stored on a machine-readable medium. In other words, one embodiment of the method according to the invention is thus a computer program that has program code for performing one of the methods described herein when the computer program is executed on a computer.
[0087] A further embodiment of the method according to the invention is thus a data carrier (or a digital storage medium or a computer-readable medium) on which the computer program for performing one of the methods described herein is recorded. The data carrier or the digital storage medium or the computer-readable medium is typically tangible and / or non-transitory.
[0088] A further embodiment of the method according to the invention is thus a data stream or a sequence of signals that represents the computer program for carrying out one of the methods described herein. The data stream or the sequence of signals can be configured, for example, to be transferred via a data communication connection, for example, via the Internet.
[0089] A further embodiment comprises a processing device, for example a computer or a programmable logic device, which is configured or adapted to carry out one of the methods described herein.
[0090] A further embodiment comprises a computer on which the computer program for performing one of the methods described herein is installed.
[0091] A further embodiment according to the invention comprises a device or system designed to transmit a computer program for performing at least one of the methods described herein to a recipient. The transmission can be electronic or optical, for example. The recipient can be, for example, a computer, a mobile device, a storage device, or a similar device. The device or system can, for example, comprise a file server for transmitting the computer program to the recipient.
[0092] In some embodiments, a programmable logic device (e.g., a field-programmable gate array, an FPGA) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field-programmable gate array may cooperate with a microprocessor to perform any of the methods described herein. In general, in some embodiments, the methods are performed by any hardware device. This may be general-purpose hardware such as a computer processor (CPU) or method-specific hardware such as an ASIC.
[0093] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. Therefore, it is intended that the invention be limited only by the scope of the following claims and not by the specific details presented in the description and explanation of the embodiments herein.
[0094] literature
[0095] Bodin Danielsson, C. & Bodin, L. (2009). Difference in satisfaction with office environment among employees in different office types. Journal of Architectural and Planning Research, 26 (3), 241-257.
[0096] Zuydervliet, R., Chanaud, R. & L’Esperance, A. (2008). Adaptive sound masking. The Journal of the Acoustical Society of America, 123 (5), 3195. https: / / doi.Org / 10.1121 / 1.2933335
[0097] Keus Van De Poll, Marijke, Carlsson, J., Marsh, J. E., Ljung, R., Odelius, J., Schlittmeier, S. J. et al. (2015). Unmasking the effects of masking on performance. The potential of multiple-voice masking in the office environment. The Journal of the Acoustical Society of America, 138 (2), 807-816. https: / / doi.Org / 10.1121 / 1.4926904
[0098] Huang, Y.-H., Robertson, M. M. & Chang, K.-l. (2004). The Role of Environmental Control on Environmental Satisfaction, Communication, and Psychological Stress. Effects of Office Ergonomics Training. ENVIRON BEHAV (Environment and behavior), 36 (5), 617- 637. https: / / doi.Org / 10.1177 / 0013916503262543
[0099] Lee, S. Y. & Brand, J. L. (2010). Can personal control over the physical environment ease distractions in office workplaces? Ergonomics, 53 (3), 324-335. https : / / d oi . org / 10.1080 / 00140130903389019
[0100] Schlittmeier, S. J. & Hellbrück, J. (2009). Background music as noise abatement in open- plan offices. A laboratory study on performance effects and subjective preferences. APPL COGNITIVE PSYCH (Applied cognitive psychology), 23 (5), 684-697. https: / / d0i.0rg / l 0.1002 / acp.1498
[0101] Hongistob, V., Varjo, J., Oliva, D., Haapakangas, A. & Benway, E. (2017). Perception of water-based masking sounds-long-term experiment in an open-plan office. FRONT PSYCHOL (Frontiers in psychology), 8. https: / / doi.org / 10.3389 / fpsyg.2017.01177
[0102] Veitch, J., Bradley, J., Legault, L, Norcross, S. & Svec, J. (2002). Masking speech in open-plan offices with simulation Ventilation noise. Noise level and spectral composition effects on acoustic satisfaction. Institute for Research in Construction. Wang, Y., Drotleff, H. A. & Li, P. (2012). Multiple maskers for speech masking in open- plan offices. The Journal of the Acoustical Society of America, 131 (4), 3481. https: / / doi.Org / 10.1121 / 1.4709135
[0103] Haapakangas, A., Kankkunen, E., Hongisto, V., Virjonen, P., Oliva, D. & Keskinen, E. (2011). Effects of five speech masking sounds on performance and acoustic satisfaction, implications for open-plan offices. ACTA ACIIST UNITED AC (Acta acustica united with Acustica), 97 (4), 641-655. https: / / doi.org / 10.3813 / AAA.918444
[0104] Chanaud, R. C. (2007). Progress in Sound Masking. Acoustics today, 3 (4), 21-26. https: / / d0i.0rg / l 0.1121 / 1.2961158
[0105] L'Esperance, A., Boudreau, A., Gariepy, F., Boudreault, L.-A. & Mackenzie, R. (2017). Adaptive Volume Control for Sound Masking Systems. How It Works and Analysis of Performance. INTER-NOISE and NOISE-CON Congress and Conference Proceedings, 254 (2), 678-686. Verfügbar unter: https: / / www.ingentaconnect.eom / content / ince / incecp / 2017 / 00000254 / 00000002 / art00083
[0106] Cohen, S. (1980). Aftereffects of stress on human performance and social behavior. A review of research and theory. Psychological bulletin, 88 (1), 82-108. https: / / doi.Org / 10.1037 / 0033-2909.88.1.82
[0107] Quick, T. L. (1990). Healthy Work. Stress, Productivity, and the Reconstruction of Working Life. National Productivity Review, 9, 475+. Verfügbar unter: https: / / link.gale. com / apps / doc / A8933314 / AONE?u=anon~d59bbb41&sid=googleScholar& xid=389d3050
[0108] Renz, T. (2019). Personalized sound masking in open offices. A trade-off between annoyance and restoration of working memory performance? Stuttgart: Fraunhofer Verl, Stuttgart.
[0109] Renz, T., Leistner, P. & Liebl, A. (eds.). (2018). A simple model to predict the cognitive performance in distracting background speech.
[0110] US 2003 / 103632 A1, Adaptive sound masking system and method, published 2003.
[0111] CN 110362789 A, published 2019. US 2015 / 194144 A1 Directional Sound Masking, published 2015.
[0112] 5
Claims
A device for generating speech masking sound, the device comprising: an analyzer (110) for analyzing each frequency-band-limited signal component of a plurality of frequency-band-limited signal components of a microphone signal during an analyzed time period to obtain information about the frequency-band-limited signal component, and a masking signal generator (120) for generating a masking signal depending on the information about the frequency-band-limited signal component of each of the plurality of frequency-band-limited signal components, the information about the frequency-band-limited signal component depending on a first sound level that was reached at least during a first time period during the analyzed time period, and the information about the frequency-band-limited signal component depending on a second sound level that was reached at least during a second time period during the analyzed time period.wherein the second time period is different from the first time period. Apparatus according to claim 1, wherein the analyzer (110) is designed to determine a microphone signal sound level difference between the first sound level and the second sound level for each signal component of the plurality of frequency-band-limited signal components of the microphone signal, and wherein the masking signal generator (120) is designed to determine the masking signal as a function of the microphone signal sound level difference of each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the microphone signal. Apparatus according to claim 2, wherein the masking signal generator (120) is designed to determine the masking signal by, for each signal component of the plurality of frequency-band-limited signal components, as a function of the microphone signal, Sound level difference of this signal component, to determine a level value for a frequency-band-limited component of the masking signal that corresponds to a frequency range of this signal component, and to perform a level adjustment of this frequency-band-limited component of the masking signal using this level value. The device according to claim 2 or 3, wherein the analyzer (110) is configured to determine an overall signal that depends on the microphone signal but is different from the microphone signal, wherein the analyzer (110) is configured to determine an error value for each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the overall signal, which error value indicates a difference between a target value for an overall signal sound level difference and a current overall signal sound level difference of the overall signal, wherein the masking signal generator (120) is configured,to determine the masking signal as a function of the error value for each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the overall signal, wherein the analyzer (110) is configured to determine the current overall signal sound level difference between a third sound level and a fourth sound level for each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the overall signal, wherein the third sound level is a sound level that was reached at least during a third time period during an analyzed time period in the frequency-band-limited signal component of the overall signal, and wherein the fourth sound level is a sound level that was reached at least during a fourth time period during the analyzed time period in the frequency-band-limited signal component of the overall signal. Device according to claim 4, wherein the analyzer (110) is configured,each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the, total signal depending on a feedback temporal section of the masking signal to this frequency band-limited signal component.
6. The device according to claim 4 or 5, wherein the analyzer (110) is configured to determine each of the plurality of frequency-band-limited signal components of the overall signal depending on an attenuation factor for this frequency-band-limited signal component of the overall signal, wherein the analyzer (110) is configured to apply the attenuation factor for this frequency-band-limited signal component to the corresponding frequency-band-limited signal component of the microphone in order to obtain an attenuated microphone signal for this frequency-band-limited signal component.
7. The device according to claim 5 and claim 6, wherein the analyzer (110) is configured to determine each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the overall signal as the sum of the feedback time portion of the masking signal to this frequency-band-limited signal component and the attenuated microphone signal to this frequency-band-limited signal component.
8. Device according to one of claims 4 to 7, wherein the masking signal generator (120) is designed to determine the masking signal as a function of a correction value for each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the overall signal, wherein the masking signal generator (120) is designed to determine the correction value for this frequency-band-limited signal component as a function of the error value for this frequency-band-limited signal component.
9. The device according to claim 8, wherein the masking signal generator (120) is further configured to determine the correction value for this frequency-band-limited signal component as a function of a temporal predecessor value of this correction value. Device according to claim 9, wherein the masking signal generator (120) is designed to determine the correction value as a function of z n = z n- + s^n(e) * ge), where z n is the correction value at a time n, where z n-1is the temporal predecessor value of this correction value at a time n-1, where e is the error value, where sgn denotes a sign function, and where g(e) denotes a function dependent on the error value e. Device according to one of the preceding claims, wherein the device is a device according to claim 2 and according to claim 4 and according to claim 9, wherein the masking signal generator (120) is designed to determine the masking signal as a function of a control value for each frequency-band-limited signal component of the plurality of frequency-band-limited signal components of the microphone signal, wherein the masking signal generator (120) is designed to determine the control value for this frequency-band-limited signal component as a function of the microphone signal-sound level difference of this frequency-band-limited signal component and as a function of the error value and the correction value of this frequency-band-limited signal component of the overall signal.Device according to claim 11, wherein the masking signal generator (120) is designed to determine the control value for this frequency-band-limited signal component by forming a sum of the microphone signal sound level difference of this frequency-band-limited signal component. and the error value and the correction value of this frequency-band-limited signal component of the overall signal. The device according to claim 11 or 12, wherein the device is a device according to claim 3, wherein the masking signal generator (120) is configured to determine the level value for a frequency-band-limited component of the masking signal as a function of the control value for this frequency-band-limited signal component and as a function of a previous level value for this frequency-band-limited component of the masking signal. A method for generating speech masking sound, the method comprising: analyzing each frequency-band-limited signal component of a plurality of frequency-band-limited signal components of a microphone signal during an analyzed period of time to obtain information about the frequency-band-limited signal component, and Generating a masking signal depending on the information about the frequency-band-limited signal component of each of the plurality of frequency-band-limited signal components, wherein the information about the frequency-band-limited signal component depends on a first sound level that was reached at least during a first time period during the analyzed time period, and wherein the information about the frequency-band-limited signal component depends on a second sound level that was reached at least during a second time period during the analyzed time period, wherein the second time period is different from the first time period. A computer program comprising program code for carrying out the method according to claim 14.