Apparatus and method for combining repetitive noise signals

By segmenting the audio signal into overlapping segments and applying time-weighted and synthesis window functions, the degradation problem of additive noise in the recorded signal is solved, thereby improving the accuracy of speaker calibration and the signal-to-noise ratio.

CN116457877BActive Publication Date: 2026-04-07FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-14
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

When measuring the transfer function of a loudspeaker in an anechoic environment or reverberation chamber, the recorded signal is degraded by additive noise such as click and bang noise, which is difficult to remove effectively with existing techniques, affecting the accuracy of the measurement and the calibration results.

Method used

The audio signal is divided into overlapping segments using segmentation blocks. An analysis window function is applied for time weighting. The weight value of each segment is calculated by determining the weight block. Finally, an overlay and summation method is used to generate the output signal using a synthesis window function to reduce the impact of noise.

Benefits of technology

It significantly improves measurement accuracy and calibration results, effectively reduces the impact of non-fixed noise, and enhances the signal-to-noise ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116457877B_ABST
    Figure CN116457877B_ABST
Patent Text Reader

Abstract

An apparatus for combining three or more audio signals is described. The apparatus includes: a segmentation block for dividing each audio signal into segments; a weight determination block configured to determine a weight value for each of the time-weighted audio signal segments; a combination block for combining the time-weighted audio signal segments of each audio signal; and a synthesis block for generating an output audio signal. A method and computer program product for combining three or more audio signals are also described.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of audio signal processing technology. Specifically, it is used to combine repetitive noise signals.

[0002] Embodiments of the present invention relate to apparatus for combining three or more audio signals. Other embodiments relate to methods for combining three or more audio signals. Other embodiments relate to using the foregoing apparatus and methods. Other embodiments relate to computer program products. Background Technology

[0003] This invention can be applied, for example, in the field of loudspeaker calibration, where repeated measurements (e.g., exponential scan measurements) are used for robust system identification. This calibration is used in modern sound systems such as soundbars and smart speakers.

[0004] When measuring the transfer function of a loudspeaker in an anechoic environment or reverberation chamber, for example, the recorded signal captured by a microphone that captures the test signal can be degraded by additive noise. In particular, non-stationary noise (such as clicks and bangs, footsteps, door closing, or background noise from fluctuations) can be problematic in practice. Reducing this noise improves the accuracy of the measurement and thus leads to better calibration results.

[0005] Transfer function measurements using exponentially scanned signals are widely used in practice because they are superior to alternative methods, such as those using maximum length sequences (MLS) as excitation signals. For practical reasons, these MLS measurements are often repeated to improve the signal-to-noise ratio. However, these repetitions cannot eliminate artifacts caused by time-varying and nonlinear distortions. These artifacts can be further reduced by using different MLS sequences.

[0006] With the introduction of improved measurements (e.g., exponentially scanned signals), repeated measurements are no longer necessary, and in fact, higher accuracy is achieved by using longer excitation signals instead of repetitions.

[0007] To address click and pop noise during recording, existing technologies employ click and pop denoising algorithms from commercial audio editors to process the recorded signal (e.g., scan signal) or use windowing methods.

[0008] This disclosure presents an improved technique for combining repetitive noise signals. Practical methods and apparatus for achieving this are described below. Summary of the Invention

[0009] Embodiments of this application relate to an apparatus for merging three or more audio signals. These audio signals are, for example, repetitive measurements of a sound system. The apparatus includes a segmentation block. The segmentation block divides each audio signal into audio signal segments. To this end, each audio signal is decomposed into multiple audio signal segments. This decomposition is performed such that each audio signal segment overlaps with adjacent audio signal segments by a predetermined percentage of its length. Of course, the first and last audio signal segments can only overlap on one side. The same segmentation is used for all audio signals such that all decomposed audio signals have corresponding segment boundaries, that is, the first, second, ..., nth audio signal segments of all audio signals have the same length, the same start time, and the same end time, respectively. The segmentation block is also configured to apply an analysis window function to each of the audio signal segments. This can be performed individually for each audio signal segment of each audio signal. Thus, each audio signal segment is transformed into a time-weighted audio signal segment.

[0010] The device also includes a weight determination block configured to determine the weight value for each of the time-weighted audio signal segments. This can also be done individually for each time-weighted audio signal segment of each audio signal.

[0011] The device also includes a combination block for combining time-weighted audio signal segments of each audio signal. This can be done individually for each audio signal. The combination is performed by calculating a weighted average of all time-weighted audio signal segments for each audio signal using determined weight values ​​for each time-weighted audio signal segment.

[0012] The device also includes a synthesis block for generating an output audio signal. This synthesis block is configured to apply a synthesis window function to a combined time-weighted audio signal segment of each audio signal and perform an overlap-add method on the corresponding results of the synthesis window function, thereby generating the output audio signal.

[0013] The proposed technique has been found to be beneficial because its performance is a significant improvement over known techniques.

[0014] According to one embodiment, the weight determination block can determine the weight value of a time-weighted audio signal segment based on the estimated noise variance value of each of the time-weighted audio signal segments, or based on the calculation of the root mean square value of the corresponding difference signal of each of the time-weighted audio signal segments.

[0015] Other alternatives are also possible.

[0016] Weight determination based on noise variance estimation has been found to be the most efficient, but the calculation of the root mean square value of the difference signal is also efficient compared to known techniques.

[0017] According to one embodiment, three or more audio signals are measurements used for speaker calibration, preferably one of the following: scan measurements, particularly preferably, exponential scan measurements; measurements using a maximum length sequence; and / or measurements using acoustic signals, particularly preferably, measurements using music.

[0018] According to one embodiment, the device decomposes an audio signal such that in each audio signal, all audio signal segments have the same length, all segments have the same percentage of overlap, and / or the same analysis window function is applied to all audio signal segments.

[0019] It has been found that each of these can improve the performance of the technology.

[0020] According to one embodiment, the overlap percentage is 50%, the analysis window function and / or the synthesis window function is one of the square roots of the cosine function and the constant overlap summation attribute window function, and / or the analysis window function and the synthesis window function are the same window function.

[0021] It has been found that each of these can improve the performance of the technology.

[0022] According to one embodiment, the product of the analytical window function and the synthetic window function satisfies the constant overlap addition property.

[0023] It has been found that this constraint is beneficial to the technology.

[0024] According to one embodiment, this device can be used for the calibration of a sound system.

[0025] Other embodiments relate to methods for combining three or more audio signals.

[0026] According to one embodiment, a method for combining three or more audio signals includes the following steps.

[0027] In the first step of this method, each audio signal is segmented into audio signal segments. These audio signals are, for example, repetitive measurements of a sound system. This segmentation involves decomposing each audio signal into multiple audio signal segments. The audio signals are decomposed such that each audio signal segment overlaps with adjacent audio signal segments by a predetermined percentage of its length. Of course, the first and last audio signal segments can only overlap on one side. The same segmentation is used for all audio signals such that all decomposed audio signals have corresponding audio signal segment boundaries; that is, the first, second, ..., nth audio signal segments of all audio signals have the same length, the same start time, and the same end time. In the first step, an analysis window function is further applied to each of the audio signal segments. This can be performed individually for each audio signal segment of each audio signal. Thus, each audio signal segment is transformed into a time-weighted audio signal segment.

[0028] In the second step of this method, a weight value is determined for each of the time-weighted audio signal segments. This can also be done individually for each time-weighted audio signal segment of each audio signal.

[0029] In the third step of this method, time-weighted audio signal segments of each audio signal are combined. This can be done individually for each audio signal. The time-weighted audio signal segments are combined by calculating a weighted average of all time-weighted audio signal segments for each audio signal using the determined weight values ​​for each time-weighted audio signal segment.

[0030] In the fourth step of this method, the output audio signal is generated by applying a synthesis window function to the combined time-weighted audio signal segments of each audio signal and performing an overlap-addition method on the corresponding results of the synthesis window function.

[0031] It has been found that the proposed technique is beneficial because its performance is significantly improved compared to known techniques.

[0032] According to one embodiment, the weight value of a time-weighted audio signal segment is determined based on either a noise variance estimate of each segment or by calculating the root mean square value of the corresponding difference signal of each segment.

[0033] Other alternatives are also possible.

[0034] It has been found that determining weight values ​​based on deterministic noise variance estimation is the most efficient method, but calculating the root mean square value of the difference signal is also efficient compared to known techniques.

[0035] According to one embodiment, three or more audio signals are measurements used for speaker calibration, preferably one of the following: scan measurements, particularly preferably, exponential scan measurements; measurements using a maximum length sequence; and / or measurements using acoustic signals, particularly preferably, measurements using music.

[0036] According to one embodiment, the audio signal is decomposed using the same length and / or the same overlap percentage for all audio signal segments, and / or the same analysis window function is applied to all audio signal segments.

[0037] It has been found that each of these can improve the performance of the technology.

[0038] According to one embodiment, a 50% overlap percentage is used to perform the decomposition step, the analysis window function and / or the synthesis window function are one of the square roots of the cosine function and the constant overlap summation attribute window function, and / or the analysis window function and the synthesis window function are the same window function.

[0039] It has been found that each of these can improve the performance of the technology.

[0040] According to one embodiment, the analytical window function and the synthetic window function are selected such that the product of the analytical window function and the synthetic window function satisfies the constant overlap addition property.

[0041] It has been found that this constraint is beneficial to the technology.

[0042] According to one embodiment, this method can be used to calibrate a sound system.

[0043] Although some aspects of this disclosure are described in conjunction with the apparatus as features, it should be understood that such descriptions can also be regarded as descriptions of corresponding method features. Similarly, although some aspects are described in conjunction with the method as features, it should be understood that such descriptions can also be regarded as descriptions of corresponding features of the device or the function of the device.

[0044] Another embodiment relates to a computer program product for implementing the above-described methods when executed on a computer or signal processor.

[0045] These methods are based on the same considerations as the aforementioned apparatus. However, it should be noted that these methods can also be supplemented by any features, functions, and details also described herein with respect to the apparatus. Furthermore, these methods can be supplemented individually and in combination by the features, functions, and details of the apparatus. Attached Figure Description

[0046] Embodiments of the present invention will now be described with reference to the accompanying drawings, in which:

[0047] Figure 1A schematic flowchart of a method according to an embodiment is shown.

[0048] Figure 2 A schematic representation of a segmented audio signal according to an embodiment is shown.

[0049] Figure 3 The schematic input and output audio signals according to an embodiment are shown.

[0050] Figure 4 A schematic diagram of a device according to an embodiment is shown, and

[0051] Figure 5 A schematic diagram of combining segments into an output signal is shown.

[0052] In the accompanying drawings, the same reference numerals denote the same elements and features. Detailed Implementation

[0053] In the following description, examples of this disclosure will be described in detail using the accompanying description. Numerous details are set forth in the description to provide a more thorough explanation of examples of this disclosure. However, it will be apparent to those skilled in the art that other examples can be implemented without these specific details. Features of the different examples described may be combined with each other unless the features of corresponding combinations are mutually exclusive or such combinations are explicitly excluded.

[0054] It should be noted that identical or similar elements with the same function may be given the same or similar reference numerals or designated as the same, and repeated descriptions of elements given the same or similar reference numerals or designated as the same are generally omitted. Descriptions of elements with the same or similar reference numerals or designated as the same are interchangeable.

[0055] In the proposed technique, three or more audio signals are combined. The audio signals represent exemplary repetitive noise signals, which may be, for example, repeated measurements of a sound system or its components. As previously mentioned, in order to measure the transfer function of such a component (e.g., a loudspeaker) in an anechoic environment or reverberation chamber, a recorded signal, for example, recorded via a microphone capturing a test signal, will be degraded by additive noise.

[0056] This audio signal represents a repeated measurement of the transfer function, i.e., the output of the sound element. In particular, non-stationary noise (such as clicks and bangs, footsteps, door closing, or background noise from fluctuations) can be detrimental to the measurement and thus negatively impact the calibration to be performed using the measurement. This calibration can be performed using continuous measurements and subsequent adjustments to the sound parameters. Other calibration methods are also possible.

[0057] Reducing the aforementioned noise improves measurement accuracy, leading to better calibration results.

[0058] Repeated measurements can be, for example, scanning measurements. Exponential scanning measurements have been found to be particularly useful. Alternative measurement techniques include measurements using the longest possible sequence and / or measurements using acoustic signals. Music, in particular, has been found to be a very subtle acoustic signal for measuring the transfer function of sound elements. Such measurements are repeated several times, with at least three repetitions required for the proposed technique.

[0059] Figure 1 A schematic flowchart of an embodiment of the proposed technology is shown. Method 100 is described in more detail below.

[0060] Method 100 begins with step 110, which is a segmentation step. Segmentation step 110 divides each audio signal 210, ..., 250 into segments.

[0061] Figure 2 Three such measurements, 210, 220, and 230, are symbolically shown; these three measurements are also referred to below as audio signals A, B, and C. As mentioned earlier, more than three measurements are possible, even if not depicted in the diagram.

[0062] Segmentation step 110 includes breaking down each audio signal into multiple audio signal segments. As an example, Figure 2 This shows that the audio signal A 210 is decomposed into segment S. A1 ... S A5 These sections are also referred to by the attached figure labels 211, ..., 215.

[0063] In substep 111, each audio signal is decomposed such that each segment of the audio signal overlaps with the adjacent segment by a predetermined percentage of the segment length. Of course, the first and last segments can only overlap on one side.

[0064] All audio signals are decomposed in the same way; that is, the same segmentation is applied to all audio signals, ensuring that all decomposed audio signals have corresponding segment boundaries. Specifically, the first, second, ..., nth segments of all audio signals have the same length, the same start time, and the same end time. The corresponding segment boundaries are... Figure 2 The values ​​are shown as vertical lines above audio signals 210, 220, 230 and audio signal segments 211, 212, 213, 214 and 215 at 0, 400, 600, 900, 1200 and 1600 ms.

[0065] Optionally, the same length is used to decompose each segment of the audio signal. If this is applied, then S A1 To S A5They will all have the same length. This is not shown in the attached diagram. Since all audio signals are similarly broken down, all segments of all audio signals have the same length. This means that if similar names are used for other audio signals B and C, then S B1 To S B5 and S C1 To S C5 Will with S A1 To S A5 They will have the same length. S B1 ... S B5 S C1 ... S C5 Not shown in the figure.

[0066] Optionally, each segment of the audio signal can have the same percentage of overlap. For ease of description, Figure 2 This has already been shown, namely 50% overlap. For example, segment S A2 It has a length of 200ms. The 50% overlap shown means that 50% of the length is with S. A1 Overlap, and 50% of the length is the same as S. A3 Overlap. In the case shown, the overlap on each side is therefore 100ms or 0.1 seconds. Other overlap percentages besides 50% can also be used. The same overlap percentage is used for all segments of all audio signals. Or the same overlap percentage is used for every nth segment of all audio signals. For example, S A1 S B1 and S C1 (abbreviated as S) X1 It can have 35% overlap, S A2 S B2 and S C2 (abbreviated as S) X2 It may have 55% overlap, etc.

[0067] In sub-step 112 of segmentation step 110, an analysis window function is applied to each audio signal segment, thereby generating time-weighted audio signal segments.

[0068] As mentioned above, since all audio signals are similarly decomposed, the analysis window function for the nth segment of each audio signal is the same. However, each segment within an audio signal can have a separate analysis window function. This means that segment S X1 It can have the same characteristics as segment S. X2 Different analysis window functions. And so on. Optionally, the analysis window function can be the same for some or all segments of an audio signal (and therefore for corresponding segments in other audio signals).

[0069] Furthermore, the analysis window function can be a cosine function. Alternatively, the analysis window function can be the square root of a constant overlap summation attribute window function, and other window functions can also be used. Constant overlap summation is also known as COLA.

[0070] The COLA window is a window function w(t) that satisfies the COLA constraints in equation (1), where T S This indicates the frame shift of a periodically applied window.

[0071] (1)

[0072] The function that satisfies this constraint is of length T. S The rectangular window is shown in equations (2) and (3).

[0073] (2)r S (t)=rect(t / T S )

[0074] (3)

[0075] Returning to the method, each segment is transformed into a time-weighted audio signal segment by segmentation step 110, and specifically by sub-step 112.

[0076] In other words, the segmentation breaks down each repeated recording into overlapping segments and applies a window function. In one embodiment, a cosine window is used as the window function. 50% overlap is a preferred embodiment. For time-aligned processing, the same segmentation is used for all repeated measurements.

[0077] In step 120, the weight value for each of the time-weighted audio signal segments is determined. This can also be done individually for each segment of each audio signal.

[0078] As an option, the weight value of a segment can be determined based on the noise variance estimate of each segment in the time-weighted audio signal.

[0079] More specifically, each segment can be modeled as x n (t)=s(t)+n n , where s(t) represents a clean signal, and n n (t) represents the additive Gaussian noise of the nth repetition. It can be assumed that the noise signals are statistically independent. Therefore, for any pair of repetitions...<i,j> For the two variance estimates involved and variance of the difference signal The calculation yields equation (4).

[0080] (4)

[0081] To determine these estimates, a system of linear equations can be constructed based on equation (5).

[0082] (5) Av = b

[0083] The following pseudocode is used to construct matrix A:

[0084]

[0085] Where N represents the number of repetitions, and M = N(N-1) / 2 represents the number of pairs. The vector b on the right-hand side of the linear equation system (5) contains the variance. And it is constructed based on the following pseudocode:

[0086]

[0087] vector It includes an unknown variance estimate. Since the linear equation system is overdetermined, the Moore-Penrose inverse A... + =(A T A) -1 A T It can be used to determine the variance estimate in the sense of minimum mean square error based on equation (6).

[0088] (6) v = A + b

[0089] Alternatively, the weight of a segment can be determined based on the root mean square value of the corresponding difference signal for each segment in the time-weighted audio signal calculation. The difference signal is determined as in the described example, except that the root is extracted and the calculation continues thereafter.

[0090] Then, method 100 proceeds to combination step 130, which combines time-weighted audio signal segments for each audio signal. This is performed individually for each audio signal. The time-weighted audio signal segments are combined by calculating a weighted average of all time-weighted audio signal segments for each audio signal using the determined weight values ​​for each time-weighted audio signal segment in sub-step 131.

[0091] According to equation (7), each repeated segment is optimally combined into a denoised segment y(t) by weighted averaging.

[0092] (7)

[0093] As discussed in one of the above options, the weight w of the current segment can be directly derived from the noise variance estimate for the current segment according to equation (8). n .

[0094] (8)

[0095] As discussed above, alternatively, the weights can be determined based on the root mean square value of the corresponding difference signal for each of the time-weighted audio signal segments.

[0096] After recombining the individual audio signals 210, ..., 250 from the modified segments, an output signal 260 is generated in generation step 140. Specifically, in substep 141, the output audio signal is generated by applying a synthesis window function to the combined segments of each audio signal. Then, in substep 142, an overlap-addition method is performed on the corresponding results of the synthesis window function, thereby generating the output audio signal.

[0097] Similar to the description of the analysis window function, since all audio signals are similarly decomposed, the synthesis window function is also similarly applied to all audio signals. This means that the synthesis window function is the same for the nth segment of each audio signal.

[0098] However, each segment within an audio signal can have a separate analysis window function, and therefore a separate synthesis window function. This means that segment S X1 It can have the same characteristics as segment S. X2 Different synthesis window functions. And so on. Optionally, the synthesis window function can be the same for some or all segments of an audio signal (and therefore for corresponding segments in other audio signals).

[0099] Furthermore, the composite window function can be a cosine function. Alternatively, the composite window function can be the square root of a constant-overlapping sum of attribute window functions, and other window functions can also be used.

[0100] Generally, in segmentation step 110, the analysis window function A is used. XY Apply to each segment S XY Above. In generation step 140, the synthesis window function SY is... XY Apply to each segment S XY Above. As mentioned above, all nth segments S X1 They will have the same analysis window function, and therefore the same synthesis window function.

[0101] However, the analysis window function A XY and the synthetic window function SY XY The same window function can be applied to some or all of these segments.

[0102] Finally, you can choose the window function to analyze window function A. XY and the synthetic window function SY XYSome or all of these conditions ensure that the product of the analytical window function and the synthetic window function satisfies the constant overlap addition property.

[0103] This can also be achieved, for example, by using a Hann or Hamming window as the analysis window instead of a synthesis window (or more precisely, by using an identity function as the synthesis window).

[0104] In other words, the final output signal 260 is generated by applying a synthesis window to the combined signal segment y(t) and performing an overlap-add method. In a preferred embodiment, a cosine window is used in the segmentation step, and the same window function is used again in the generation step to obtain constant overlap-add properties.

[0105] Figure 3 An example of an embodiment of the proposed technique with five repetitions (i.e., an audio signal, which may be, for example, an analog recording) is shown. For example, the audio signal includes: for example, non-fixed signal degradation, as shown in inputs 1 to 4210, ..., 240, and different noise levels, as shown in input 5250. Output signal 260 is shown as the result. Each signal is shown with the x-axis indicating time in seconds and the y-axis indicating x(t).

[0106] Figure 4 A device 400 for combining three or more audio signals 210, ..., 250 is shown. These audio signals 210, ..., 250 are, for example, repetitive measurements of a sound system. The device includes a segmentation block 410. The segmentation block 410 divides or decomposes each audio signal 210, ..., 250 into multiple segments 211, ..., 215. This decomposition is performed such that each segment overlaps with the adjacent segment by a predetermined percentage of segment length. Of course, the first and last segments can only overlap on one side. The same segmentation is used for all audio signals such that all decomposed audio signals have corresponding segment boundaries, that is, the first, second, ..., nth segments of all audio signals have the same length, the same start time, and the same end time, respectively. The segmentation block is also configured to apply an analysis window function to each of the audio signal segments. This can be performed individually for each segment of each audio signal. Thus, each segment is transformed into a time-weighted audio signal segment.

[0107] The device also includes a weight determination block 420 configured to determine a weight value for each of the time-weighted audio signal segments. This can also be done individually for each segment of each audio signal.

[0108] The device also includes a combination block 430 for combining time-weighted audio signal segments of each audio signal. This can be done individually for each audio signal. The combination is performed by calculating a weighted average of all time-weighted audio signal segments of each audio signal using determined weight values ​​for each time-weighted audio signal segment.

[0109] The device also includes a synthesis block 440 for generating an output audio signal. This synthesis block is configured to apply a synthesis window function to each combined segment of the audio signal and perform an overlap-add method on the corresponding results of the synthesis window function, thereby generating the output audio signal.

[0110] Figure 5 An example of the effect of this method on audio signal 510 is shown. The first audio signal 510 is decomposed (sub-step 111 above) into segments starting with k. These segments are referred to as 511, ..., 514, and these segments overlap with 50% overlap as schematically shown. Then, an analysis window function is applied to each of the audio signal segments (sub-step 112 above) in 520, ..., 550 to generate time-weighted audio signal segments 521, ..., 524. These time-weighted audio signal segments 521, ..., 524 are then combined again using weights that were determined simultaneously with or before the combination (step 120 above) to form the processed audio signal 560.

[0111] If each audio signal has been processed in this way, then the processed audio signals are recombined (step 130 above). Figure 5 (not shown in the image) to form the output signal.

[0112] The methods and apparatus described above can be used to calibrate sound systems.

[0113] In summary, the proposed technique employs repetitive audio signals (such as exponential scan measurements repeated several times (at least 3 times)) and, as an example, continuously estimates the short-term variance of the additive noise for each repetition. Then, using time-varying variance estimation, repeated measurements are combined using a weighted average in the sense of minimum mean square error.

[0114] Advantageously, if one (or more) of the repeating audio signals (i.e., scan recordings) exhibits a significantly larger noise variance at a given time than other recordings, then a significantly smaller weight will be applied to that signal segment. Therefore, the proposed method can handle non-stationary noise well. Figure 3 This is illustrated.

[0115] Compared to the proposed technique, conventional methods do not handle non-stationary noise well. If the recorded scan contains some unwanted background noise, the measurement must be repeated.

[0116] In summary, the embodiments described herein may optionally be supplemented by any important points or aspects described herein. However, it should be noted that the important points and aspects described herein may be used alone or in combination, and may be incorporated, alone or in combination, into any embodiment described herein.

[0117] Although some aspects have been described in the context of the apparatus, it will be clear that these aspects also represent a description of the corresponding method, wherein the apparatus or a portion thereof corresponds to a method step or a feature of the method step. Similarly, aspects described in the context of method steps also represent a description of the corresponding apparatus or a portion thereof, or a feature of the corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware devices (such as microprocessors, programmable computers, or electronic circuits). In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0118] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or software. Implementation can be performed using a digital storage medium (e.g., floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) on which electronically readable control signals are stored, which cooperate with (or are capable of cooperating with) a programmable computer system to perform the corresponding methods. Therefore, the digital storage medium can be computer-readable.

[0119] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0120] Typically, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product is run on a computer. This program code may, for example, be stored on a machine-readable medium.

[0121] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.

[0122] In other words, embodiments of the method of the present invention are therefore computer programs having program code for performing one of the methods described herein when the computer program is run on a computer.

[0123] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) including a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory.

[0124] Therefore, another embodiment of the method of the present invention represents a data stream or signal sequence of a computer program used to perform one of the methods described herein. This data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).

[0125] Another embodiment includes a processing means, such as a computer or a programmable logic device, which is configured or adapted to perform one of the methods described herein.

[0126] Another embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.

[0127] Another embodiment of the invention includes an apparatus or system configured to transmit a computer program to a recipient (e.g., electronically or optically), the computer program being used to perform one of the methods described herein. The recipient may be, for example, a computer, mobile device, storage device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the recipient.

[0128] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.

[0129] The apparatus described herein can be implemented using hardware devices, a computer, or a combination of hardware devices and a computer.

[0130] The apparatus described herein, or any component thereof, may be implemented, at least in part, in hardware and / or software.

[0131] The methods described herein can be performed using hardware devices, computers, or a combination of hardware devices and computers.

[0132] The methods described herein, or any part thereof, may be performed at least in part by hardware and / or software.

[0133] The above embodiments are merely illustrative of the principles of the invention. It should be understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, it is intended to be limited only by the scope of the appended claims, and not by the specific details given by way of the description and explanation of the embodiments herein.

Claims

1. An apparatus for combining three or more audio signals, the apparatus comprising: A segmentation block, used to segment each audio signal, is configured to: decompose each audio signal into multiple audio signal segments, each audio signal segment overlapping with adjacent audio signal segments by a predetermined percentage of its length, wherein all decomposed audio signals have corresponding audio signal segment boundaries, such that the first, second, ..., nth audio signal segments of all audio signals have the same length, the same start time, and the same end time; and apply an analysis window function to each of the audio signal segments to generate time-weighted audio signal segments. The weight determination block is configured to determine the weight value for each of the time-weighted audio signal segments. A combining block is configured to combine the time-weighted audio signal segments of each audio signal, the combining block being configured to calculate a weighted average of all time-weighted audio signal segments of each audio signal using a determined weight value for each time-weighted audio signal segment. A synthesis block is used to generate an output audio signal. The synthesis block is configured to apply a synthesis window function to a combined time-weighted audio signal segment of each audio signal and perform an overlap-add method on the corresponding result of the synthesis window function.

2. The apparatus according to claim 1, wherein, The weight determination block is configured to determine the weight values ​​of the time-weighted audio signal segments based on the following: The determination of the noise variance estimate for each of the time-weighted audio signal segments, or Calculation of the root mean square value of the corresponding difference signal for each of the time-weighted audio signal segments.

3. The apparatus according to claim 1, wherein, The three or more audio signals are measurements used for speaker calibration, and the measurements are one of the following: scan measurements, including exponential scan measurements; measurements using a maximum length sequence; and measurements using acoustic signals, including measurements using music.

4. The apparatus according to claim 1, wherein, For each audio signal, all audio signal segments have the same length, all audio signal segments have the same overlap percentage, and / or the same analysis window function is applied to all audio signal segments.

5. The apparatus according to claim 1, wherein The segmentation blocks are configured to use a 50% overlap percentage to decompose each audio signal. The analysis window function and / or the synthesis window function is one of the square roots of a cosine function or any window function with constant overlap and addition properties, and / or The analysis window function and the synthesis window function are the same window function.

6. The apparatus according to claim 1, wherein, The product of the analysis window function and the synthesis window function satisfies the constant overlap and addition property.

7. The apparatus according to claim 1, used for sound system calibration.

8. A method for combining three or more audio signals, comprising: Segment each audio signal, including: Each audio signal is decomposed into multiple audio signal segments, each segment overlapping with adjacent segments by a predetermined percentage of its length. All decomposed audio signals have corresponding segment boundaries, such that the first, second, ..., nth audio signal segments of all audio signals have the same length, the same start time, and the same end time. The analysis window function is applied to each of the audio signal segments to generate a time-weighted audio signal segment. Determine the weight value for each of the time-weighted audio signal segments. The time-weighted audio signal segments that combine each audio signal include: The weighted average of all time-weighted audio signal segments for each audio signal is calculated using the determined weight values ​​for each time-weighted audio signal segment, and Generate the output audio signal, including: The synthesis window function is applied to the combined time-weighted audio signal segments of each audio signal, and The overlapping addition method is performed on the corresponding result of the synthesis window function.

9. The method according to claim 8, wherein, The weight values ​​of the time-weighted audio signal segments are determined based on the following: Determine the noise variance estimate for each of the time-weighted audio signal segments, or Calculate the root mean square value of the corresponding difference signal for each of the time-weighted audio signal segments.

10. The method according to claim 8, wherein, The three or more audio signals are measurements used for speaker calibration, and the measurements are one of the following: scan measurements, including exponential scan measurements; measurements using a maximum length sequence; and / or measurements using acoustic signals, including measurements using music.

11. The method according to claim 8, wherein, For each audio signal, the decomposition steps are performed using the same length and / or the same overlap percentage for all audio signal segments, and / or the same analysis window function is applied to all audio signal segments.

12. The method of claim 8, wherein Use a 50% overlap percentage to perform the decomposition step. The analysis window function and / or the synthesis window function is one of the square roots of a cosine function or any window function with constant overlap and addition properties, and / or The analysis window function and the synthesis window function are the same window function.

13. The method according to claim 8, wherein, The product of the analysis window function and the synthesis window function satisfies the constant overlap and addition property.

14. The method of claim 8, used for calibrating a sound system.

15. A computer program product for implementing the method of claim 8 when executed on a computer or signal processor.