A method and system for processing a multi-channel speech signal containing wind noise

By performing frame division, time-frequency transformation, and position grouping detection on multi-channel signals from multi-microphone devices, and combining channel energy and spectral distribution to select the optimal channel output, the problem of large wind noise differences in multi-microphone devices is solved, and audio processing quality and signal stability are improved.

CN116580722BActive Publication Date: 2025-12-05GOERTEK INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310511864.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-05
Publication Date
2025-12-05
Estimated Expiration
2043-05-05

AI Technical Summary

Technical Problem

In multi-microphone devices, the wind noise content in the voice signals collected by each microphone channel varies, resulting in significant differences in wind noise. Existing technologies struggle to effectively select the optimal channel with the least wind noise for output in real time, thus affecting audio processing quality.

Method used

By performing frame division and time-frequency transformation preprocessing on multi-channel signals input from multiple microphones, wind noise detection is performed by grouping according to microphone positions. Combining the time-domain energy magnitude and frequency-domain spectrum distribution of the channels, the channel signal with the lowest energy or the highest spectral centroid is selected frame by frame from all channel signals for output. Weighted smoothing is then used to reduce spectral abrupt changes caused by channel switching.

Benefits of technology

It enables real-time and accurate identification of the optimal channel with the least wind noise from multiple channels, improving the quality of audio processing, reducing spectral abrupt changes caused by channel switching, and ensuring the stability and clarity of the output signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116580722B_ABST
    Figure CN116580722B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-channel wind noise containing speech signal processing method and system, comprising: multi-channel signal is respectively framed and time-frequency conversion preprocessing;According to microphone position, the multi-channel signal is grouped, and each channel group includes at least 2 channel signals;With channel group as unit, wind noise detection is respectively carried out to each channel group to determine whether it contains wind noise;When determining that all channel groups or part of channel groups do not contain wind noise, the channel signal without wind noise is output after frame beam fusion;When determining that all channel groups contain wind noise, the channel signal with minimum energy or highest spectral barycenter is selected from all channel signals frame by frame to output in combination with the time energy size and frequency domain spectrum distribution of channel.The application can accurately identify the optimal channel with least wind noise from multiple channels in real time and effectively output.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing, in particular to a multi-channel wind noise containing speech signal processing method and system. BACKGROUND

[0002] In the process of collecting speech signals, microphones (MICs) are inevitably affected by the surrounding environment, including wind noise generated in the wind environment. For multi-MIC electronic devices, since the environmental wind has a certain blowing direction, the wind noise contained in the speech signals collected by each MIC channel is not the same.

[0003] Under normal circumstances, the differences in speech signals picked up by each MIC are not large, but because the wind has a direction, the wind noise picked up by each channel MIC can have a large difference, and the energy of the entire channel with large wind noise is also large. In order to improve the audio processing quality of multi-MIC electronic devices, it is necessary to select the channel with the least wind noise content from multiple MIC channels in real time for subsequent speech processing. SUMMARY

[0004] The embodiments of the present application provide a multi-channel wind noise containing speech signal processing method and system, which can accurately identify the optimal channel with the least wind noise content from multiple MIC channels in real time and effectively output.

[0005] According to a first aspect of the present application, a multi-channel wind noise containing speech signal processing method is provided, the method comprising:

[0006] frame and time-frequency transformation preprocessing are performed on the multi-channel signals input by multiple microphones in real time;

[0007] The multi-channel signals are grouped according to microphone positions, and each channel group includes at least two channel signals;

[0008] Wind noise detection is performed on each channel group to determine whether wind noise is contained;

[0009] When it is determined that all channel groups or part of the channel groups do not contain wind noise, the speech signals of all channels in the channel group without wind noise are output after beam fusion frame by frame;

[0010] When it is determined that all channel groups contain wind noise, the channel signal with the smallest energy or the highest spectral barycenter is selected from all channel signals frame by frame according to the time energy size and frequency spectrum distribution of the channel to output.

[0011] In a preferred embodiment, the channel signal with the smallest energy or the highest spectral barycenter is selected from all channel signals frame by frame according to the time energy size and frequency spectrum distribution of the channel to output, comprising:

[0012] selecting two channels with minimum energy and second minimum energy from all the channel signals frame by frame, dividing the energy of the two channels, if the quotient is in a preset range close to 1, calculating the cumulative spectrum average gravity factor of each channel, and selecting the channel signal corresponding to the cumulative spectrum average gravity factor with larger value to output, otherwise selecting the channel signal with minimum energy to output.

[0013] As an improvement of the above scheme, the method further comprises:

[0014] In the process of outputting the channel signal, the channel number selected in the current frame is compared with the channel number output in the last frame, if equal, the channel signal of the channel number selected in the current frame is directly outputted;

[0015] If not equal, the channel signal of the channel number output in the last frame is continuously outputted until a preset number of frames is reached.

[0016] As a further improvement of the above scheme, the method further comprises:

[0017] After the channel switching, the channel signal of the channel number in the last frame is weighted and smoothed frame by frame with the channel signal of the channel number in the current frame, to obtain a weighted and smoothed speech signal, wherein the weight of the channel signal of the channel number in the current frame increases frame by frame;

[0018] The weighted and smoothed speech signal is outputted frame by frame until a preset number of smoothing frames is reached.

[0019] In a preferred embodiment, the calculation process of the cumulative spectrum average gravity factor comprises:

[0020] The single-frame spectrum gravity factor is obtained according to the ratio of the frame amplitude frequency product and the frame amplitude value, wherein the frame amplitude frequency product is obtained by summing the product of the amplitude value of each frequency point and the corresponding frequency in each frame, and the frame amplitude value is obtained by summing the amplitude value of all frequency points in each frame;

[0021] The cumulative spectrum average gravity factor is obtained by averaging the sum of all single-frame spectrum gravity factors in the current cumulative frame.

[0022] According to the second aspect of the present application, a multi-channel wind noise containing speech signal processing system is provided, the system comprises:

[0023] A preprocessing module is configured to perform frame division and time-frequency transformation preprocessing on the multi-channel signals input by multiple microphones in real time respectively;

[0024] A position grouping module is configured to group the multi-channel signals according to the positions of the microphones, each channel group comprising at least two channel signals;

[0025] a wind noise detection module, configured to respectively perform wind noise detection on each channel group in a unit of channel group to determine whether the wind noise is contained;

[0026] a beam fusion module, configured to output all channel signals in a channel group without wind noise after frame-by-frame beam fusion of the channel signals when it is determined that all channel groups or only part of the channel groups do not contain wind noise;

[0027] a channel selection module, configured to select a channel signal with minimum energy or highest spectral barycenter from all channel signals frame by frame when it is determined that all channel groups contain wind noise, in combination with time domain energy size and frequency domain spectral distribution of the channel.

[0028] In a preferred embodiment, the channel selection module comprises:

[0029] an energy screening unit, configured to select two channels with minimum energy and sub-minimum energy from all channel signals frame by frame when it is determined that all channel groups contain wind noise.

[0030] an energy comparison unit, configured to divide the energy of the two channels selected by the energy screening unit to determine whether a quotient obtained by the division is within a preset range close to 1.

[0031] a spectral barycenter factor calculation unit, configured to calculate respective cumulative spectral average barycenter factors of the two channels.

[0032] a first channel selection unit, configured to select a channel signal corresponding to a cumulative spectral average barycenter factor with a larger value to output when the energy comparison unit determines that the quotient obtained by dividing the energy of the two channels is within the preset range close to 1.

[0033] a second channel selection unit, configured to select a channel signal with minimum energy to output when the energy comparison unit determines that the quotient obtained by dividing the energy of the two channels is not within the preset range close to 1.

[0034] As an improvement of the above scheme, the system further comprises:

[0035] a channel retention module, configured to compare a channel number selected in a current frame with a channel number output in a previous frame during output of the channel signal, and directly output a channel signal of the channel number selected in the current frame if the channel numbers are equal, or continue to output the channel signal of the channel number output in the previous frame until a preset number of frames is reached if the channel numbers are not equal.

[0036] As an improvement of the above scheme, the system further comprises:

[0037] The smooth switching module is configured to, after channel switching, weight and smooth the channel signal of the previous frame channel number and the channel signal of the current frame channel number frame by frame to obtain a weighted and smoothed voice signal, wherein the weight of the channel signal of the current frame channel number increases frame by frame; and output the weighted and smoothed voice signal frame by frame until a preset number of smoothing frames is reached.

[0038] In a preferred embodiment, the spectral centroid factor calculation unit calculates the cumulative spectral average centroid factor in the following way:

[0039] The single-frame spectral centroid factor is obtained according to the ratio of the frame amplitude frequency product to the frame amplitude value, wherein the frame amplitude frequency product is obtained by summing the products of the amplitudes of all frequency points in each frame and the corresponding frequencies, and the frame amplitude value is obtained by summing the amplitudes of all frequency points in each frame.

[0040] The cumulative spectral average centroid factor is obtained by averaging the sum of all single-frame spectral centroid factors in the current cumulative frame.

[0041] According to a third aspect of the present application, an electronic device is provided, comprising a plurality of microphones, a memory and a processor, wherein the plurality of microphones are configured to collect voice signals of the surrounding environment in real time to obtain multi-channel signals respectively;

[0042] The memory stores a computer program, and the computer program is loaded and executed by the processor to implement the multi-channel wind noise containing voice signal processing method described above.

[0043] According to a fourth aspect of the present application, a computer readable storage medium is provided, which stores one or more computer programs, and the one or more computer programs, when executed by a processor, implement the multi-channel wind noise containing voice signal processing method described above.

[0044] The present application has the following advantages:

[0045] The embodiment of the present application utilizes the characteristics that the wind noise contained in the different channel voice signals picked up by each MIC is different in size, comprehensively considers the time domain characteristics and frequency domain characteristics of the audio signal, when it is determined that the voice signals of all channels contain wind noise, the channel signal with the minimum energy or the highest spectral barycenter is selected from all channel signals frame by frame for output in combination with the time domain energy size and the frequency domain spectrum distribution of the channel, so that the characteristics that the channel with large wind noise also has large channel energy are utilized, the channel signal with the minimum energy is selected for output through the time domain energy decision strategy, at the same time, in order to overcome the possible local optimum of the single energy minimum decision strategy, the characteristics that the wind noise energy is mostly concentrated in the low frequency are utilized, the energy frequency distribution is identified through the frequency domain spectrum, the channel signal with the highest spectral barycenter is selected for output, so that the problem of misselecting the suboptimal channel in the optimal channel selection process is solved, so that the optimal channel with the least wind noise content can be accurately identified from multiple channels in real time and effectively output. BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art according to these drawings. In the drawings:

[0047] Figure 1 A flowchart of a multi-channel wind noise containing voice signal processing method of an embodiment of the present application is shown;

[0048] Figure 2 A frame division schematic diagram is shown;

[0049] Figures 3(a) to 3(d) Time domain waveform diagrams and corresponding spectrum diagrams of four channel signals when the wind blows from the front in an embodiment of the present application are shown, wherein Fig. 3(a) corresponds to L1 channel, Fig. 3(b) corresponds to R1 channel, Fig. 3(c) corresponds to R2 channel, and Fig. 3(d) corresponds to L2 channel;

[0050] Figure 4 Time domain waveform diagrams and corresponding spectrum diagrams of four channel signals of Fig. 3 when the wind blows from the front according to the signal energy minimum decision strategy simulation output are shown;

[0051] Figures 5(a) to 5(d) Time domain waveform diagrams and corresponding spectrum diagrams of four channel signals when the wind blows from the left in an embodiment of the present application are shown, wherein Fig. 5(a) corresponds to L1 channel, Fig. 5(b) corresponds to R1 channel, Fig. 5(c) corresponds to R2 channel, and Fig. 5(d) corresponds to L2 channel;

[0052] Figure 6Figures showing the time-domain waveform diagram and corresponding frequency spectrum diagram of the output of the 4-channel signal of Fig. 5 according to the energy-minimum simulation decision strategy when the wind blows from the left side;

[0053] Figure 7 Figures showing the gc K curve of the Rl channel corresponding to Fig. 5(b) and the gc

[0054] Figure 8 Figures showing the frequency spectrum local comparison diagram of the output signal of the channel switching of an embodiment of the present application before and after the improvement of channel maintenance and smooth switching;

[0055] Figure 9 Figures showing the structure diagram of the processing system of the multi-channel wind-noise-containing speech signal of an embodiment of the present application;

[0056] Figure 10 Figures showing the structure diagram of the channel selection module of an embodiment of the present application;

[0057] Figure 11 Figures showing the structure diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0058] Embodiments of the present application will be described in more detail with reference to the drawings. These embodiments are provided so that the present application can be more thoroughly understood and so that the scope of the present application can be completely conveyed to those skilled in the art. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein.

[0059] Figure 1 Figures showing the flow diagram of the processing method of the multi-channel wind-noise-containing speech signal of an embodiment of the present application. Referring to Figure 1 the method of the present application includes the following steps S110 to S150:

[0060] Step S110, performing frame division and time-frequency transformation preprocessing on the multi-channel signal input in real time by a plurality of microphones.

[0061] The preprocessing performed by the present step S110 on the multi-channel signal includes: frame division in the time domain, referring to Figure 2 , Figure 2 a frame division diagram is shown, the length of each frame and the length of overlap (or frame shift) can be set as needed, generally the length of each frame is about 7.5ms to 15ms, and windowed short-time Fourier transformation is performed on the time-domain data of each frame to obtain the frequency-domain signal of each frame.

[0062] Step S120, grouping the multi-channel signals according to microphone positions, each channel group including at least two channel signals.

[0063] The multiple microphones on the electronic device will have different sound signals due to different spatial positions. In the presence of wind, each microphone will input a channel signal in real time, which includes both voice signals and wind noise. However, in general cases, the energy difference of voice signals in each channel will not be large, but the wind noise energy in each channel may be quite different because of the wind direction.

[0064] Grouping the multi-channel signals according to microphone positions, each channel group including at least two channel signals, this grouping method can divide the microphones close to each other into the same channel group, ensuring that the wind noise picked up by each channel in the same channel group is close in size.

[0065] In the following description, for the convenience of simulation and writing, four channels are taken as an example, which can be divided into L and R two groups. For the convenience of description, the numbers of the four channels are L1, L2 and R1, R2 respectively.

[0066] It should be noted that for other multi-channel numbers, such as three channels, they can be divided into left and right groups by sharing the MIC in the middle position. For example, six channels can be divided into three groups, each group including two channels, or divided into two groups, each group including three channels. As for how to group, it needs to be adjusted in combination with product hardware design and layout.

[0067] Step S130, performing wind noise detection on each channel group in units of channel groups to determine whether it contains wind noise.

[0068] The wind noise detection can use any existing method, such as a two-channel wind noise detection method, which can realize wind noise detection based on the commonly used chi-square algorithm. The specific wind noise detection method used by the present application is not limited, so it is not described in detail.

[0069] In the above four-channel example, for the L and R two channel groups, wind noise detection is performed on each channel group in units of channel groups to determine whether it contains wind noise. According to the detection results, different values can be returned in the program, for example, 0 when there is no wind on both L and R sides, 1 when there is wind on the L side, 2 when there is wind on the R side, and 3 when there is wind on both L and R sides.

[0070] Step S140, when it is determined that all channel groups or part of the channel groups do not contain wind noise, performing beam fusion on all channel signals in the channel group without wind noise frame by frame and then outputting.

[0071] Continuing the above 4-channel example, when there is no wind, the return value is 0, then the 4-channel signals are frame by frame beam fused and output; when there is wind but the wind direction is relatively fixed, it is detected that only one side channel group contains wind noise, and the other side channel group does not contain wind noise, that is, the return value is 1 or 2, then only the L or R side two-channel signal which does not contain wind noise is frame by frame beam fused and output.

[0072] By frame by frame beam fusing all channel signals in the channel group which does not contain wind noise and then outputting, the output signal can be ensured not to contain wind noise, and the output signal energy is large.

[0073] Step S150, when it is determined that all channel groups contain wind noise, the channel signal with the minimum energy or the highest spectral center of gravity is selected from all channel signals frame by frame according to the time domain energy size and frequency domain spectrum distribution of the channel to output.

[0074] Continuing the above 4-channel example, when there is wind and the wind direction is not fixed, it is detected that both sides of the channel group contain wind noise, the return value is 3, and considering that the channel energy of the channel containing wind noise is large, the expected processing at this time is to select a channel signal with the minimum energy in the 4 channels by the time domain energy decision strategy. The algorithm formula used is as follows:

[0075]

[0076]

[0077] In the above two formulas, S ch k represents the input signal of the ch channel and the k frame, S k,out N is the number of sampling points per frame, and i is the sampling sequence number per frame.

[0078] The following will be described in conjunction with specific 4-channel signals.

[0079] Figures 3(a) to 3(d) The time domain waveform diagram and the corresponding spectrum diagram of the 4-channel signals when the wind blows from the front are shown, wherein Fig. 3(a) corresponds to the L1 channel, Fig. 3(b) corresponds to the R1 channel, Fig. 3(c) corresponds to the R2 channel, and Fig. 3(d) corresponds to the L2 channel.

[0080] First, the whole is described as follows Figures 3(a) to 3(d) Each figure includes two parts, the upper half is a time domain waveform diagram, and the vertical coordinate represents the signal amplitude; the lower half is a spectrum diagram, and the vertical coordinate represents the signal frequency; the upper and lower parts share the horizontal coordinate, and the horizontal coordinate represents the time t. The subsequent similar figure description is the same as above, and will not be repeated here.

[0081] Considering that the voice signals received by the respective channels are not much different, but the wind noise is different, the comparison Figures 3(a) to 3(d) It can be seen that the wind noise picked up by the four channels is different when the wind blows from the front. The optimal channel with the lowest wind noise content is usually selected according to the minimum energy, and the simulation output result is shown in Figure 4 . Figure 4 The time-domain waveform diagram and the corresponding frequency spectrum diagram of the simulation output of the four-channel signal of Figure 3 according to the minimum energy decision strategy when the wind blows from the front are shown.

[0082] However, since the coordinates of the four MICs are different, and the wind direction changes at any time, the amplitude of the received noisy voice signal is different, and the amplitude of the channel signal without wind noise is not necessarily smaller than the amplitude of the channel signal with wind noise at all times, so the suboptimal channel may be selected by mistake according to the single minimum energy decision strategy.

[0083] Figures 5(a) to 5(d) The time-domain waveform diagram and the corresponding frequency spectrum diagram of the four-channel signal of an embodiment of the present application when the wind blows from the left are shown, wherein Figure 5(a) corresponds to the L1 channel, Figure 5(b) corresponds to the R1 channel, Figure 5(c) corresponds to the R2 channel, and Figure 5(d) corresponds to the L2 channel. Comparing Figures 5(a) to 5(d) It can be seen that when the wind blows from the left, the left channel group (L1 channel and L2 channel) has more wind noise content, and the right channel group (R1 channel and R2 channel) has less wind noise content, and the optimal channel is the R1 channel corresponding to Figure 5(b) in the entire time period.

[0084] Figure 6 The time-domain waveform diagram and the corresponding frequency spectrum diagram of the simulation output of the four-channel signal of Figure 5 according to the minimum energy decision strategy when the wind blows from the left are shown. However, it can be seen from the hatched part in Figure 6 that there are some vertical lines in the frequency spectrum, indicating that the channel output according to the minimum energy decision strategy is selected between the R1 channel and the R2 channel, that is, not all frames are the optimal R1 channel, but some frames are the suboptimal R2 channel.

[0085] To avoid the selection of a suboptimal channel according to the single minimum energy decision strategy in the optimal channel selection process, considering the feature that the wind noise energy is mostly concentrated in the low frequency, the channel signal with the highest spectrum center of gravity is selected for output in this step S150, so as to solve the problem of the selection of a suboptimal channel in the optimal channel selection process, thereby accurately identifying the optimal channel with the least wind noise content from the multiple channels in real time and effectively.

[0086] In a preferred embodiment, the "combining the time-domain energy size and the frequency-domain spectral distribution of the channels to select the channel signal with the minimum energy or the maximum spectral center from the all channel signals frame by frame for output" of the present step S150 specifically includes:

[0087] The two channels with the minimum energy and the second minimum energy are selected from the all channel signals frame by frame, the energy of the two channels is divided, if the quotient is in a preset range close to 1, the cumulative spectral average center factors of the two channels are calculated, and the channel signal corresponding to the cumulative spectral average center factor with the larger value is selected for output, otherwise the channel signal with the minimum energy is selected for output.

[0088] In the case that all the channel groups contain wind noise, the present preferred embodiment first selects the two channels with the minimum energy and the second minimum energy from the all channel signals frame by frame according to the time-domain energy decision strategy, and excludes other channels with larger energy, i.e. containing more wind noise, so as to reduce the complexity, narrow the channel selection range, and reduce the subsequent calculation amount without substantially affecting the accuracy of the optimal channel selection.

[0089] In order to obtain the frequency-domain spectral distribution of the channel, the present preferred embodiment introduces a single-channel cumulative spectral average center factor, and the calculation process is as follows:

[0090] The single-frame spectral center factor is obtained according to the ratio of the frame amplitude frequency product to the frame amplitude value, wherein the frame amplitude frequency product is obtained by summing the products of the amplitudes of all frequency points and the corresponding frequencies of each frame, and the frame amplitude value is obtained by summing the amplitudes of all frequency points of each frame.

[0091] The sum of all single-frame spectral center factors in the current cumulative frame is averaged to obtain the cumulative spectral average center factor.

[0092] The above calculation process is described by the following formula:

[0093] Step 1: Calculate the single-frame spectral center factor:

[0094]

[0095] In the above formula 3, k represents the kth frame, F k (i) represents the amplitude of the ith frequency point after the Fourier transform of the kth frame, M is the number of short-time Fourier transform points, f i is the frequency corresponding to the ith frequency point, f s is the sampling frequency.

[0096] In addition, considering the symmetry of the spectrum, only the first half of the frequency points, i.e. M / 2+1 frequency points, need to be calculated.

[0097] Second step: calculate the cumulative spectrum average gravity center factor:

[0098]

[0099] By averaging the sum of all single-frame spectrum gravity center factors in the current cumulative frame, all single-frame spectrum gravity center factors in the current cumulative frame are normalized, which is equivalent to average smoothing in time domain, avoiding poor robustness of single-frame calculation.

[0100] For the four channel signals shown in Figure 5 when the wind blows from the left side, the L1 channel and the L2 channel of the left channel group can be easily or intuitively excluded according to the energy minimum decision strategy, but for the R1 channel and the R2 channel of the right channel group, the cumulative spectrum average gravity center factor gc K .

[0101] Figure 7 The respective gc K curves of the R1 channel corresponding to Figure 5(b) and the R2 channel corresponding to Figure 5(c) are shown, where the horizontal coordinate represents time t and the vertical coordinate represents signal frequency. The upper gc K curve corresponds to the R1 channel, and the lower gc K curve corresponds to the R2 channel. It can be seen from Figure 7 that the gc K value of the R1 channel with less wind noise is large, indicating that the cumulative spectrum average gravity center is high, and the gc K value of the R2 channel with more wind noise is small, indicating that the cumulative spectrum average gravity center is low, so by comparing the size of the gc K curve, that is, by comparing the height of the cumulative spectrum average gravity center, the respective wind noise conditions of the R1 channel and the R2 channel can be distinguished more obviously.

[0102] In the preferred embodiment, when a channel signal is selected for output according to the energy minimum decision strategy frame by frame, the two channels with the minimum energy and the second minimum energy are first selected from all channels, denoted as index1 and index2, and the energy energy(index1) and energy(index2) of the two channels are divided. If the quotient energy(index1) / energy(index2) is in a preset range close to 1, for example, in the range of 0.8-1.2, it means that the energies of the two channels are close, and to avoid selecting a suboptimal channel, the gc K value of the two channels is then calculated, and the gc KThe channel signal corresponding to the larger value is selected, and the channel signal containing more low-frequency energy, that is, more wind noise, is excluded; if the quotient energy(index1) / energy(index2) is not in the preset range close to 1, it indicates that the energy of the two channels is quite different, and at this time, the channel with the minimum energy is directly output, and the sub-optimal channel is not selected by mistake.

[0103] In summary, the method for processing a multi-channel wind noise containing voice signal provided by the application first performs frame division and time-frequency transformation preprocessing on the multi-channel signals input in real time by multiple microphones, and groups the multi-channel signals according to the microphone positions, so that wind noise detection can be performed on each channel group to determine whether wind noise is contained; in the case where it is determined that all channel groups or part of the channel groups do not contain wind noise, all channel signals in the channel group without wind noise are output after being frame by frame beam fused, so that the output signal is ensured to not contain wind noise and the energy of the output signal is large; in the case where it is determined that all channel groups contain wind noise, considering the actual situation that the amplitude of the channel signal without wind noise is not necessarily smaller than the amplitude of the channel signal with wind noise in real time, to avoid selecting a sub-optimal channel by mistake, the channel signal with the minimum energy or the highest spectral barycenter is selected from all channel signals frame by frame to be output, in combination with the time domain energy size and the frequency domain spectral distribution of the channel.

[0104] It can be seen that the method of the application utilizes the different wind noise content in different channel voice signals picked up by each MIC, comprehensively considers the time domain characteristics and frequency domain characteristics of the audio signal, utilizes the characteristics that the channel energy of the channel with more wind noise is also large, selects the channel signal with the minimum energy for output through the time domain energy decision strategy, and at the same time, in order to overcome the possible local optimization of the single energy minimum decision strategy, the characteristics that wind noise energy is mostly concentrated in the low frequency are utilized, the energy frequency distribution is identified through the frequency domain spectrum, the channel signal with the highest spectral barycenter is selected for output, and the problem of selecting a sub-optimal channel by mistake in the optimal channel selection process is solved, so that the optimal channel with the least wind noise content can be accurately identified from multiple channels in real time and effectively output.

[0105] In the process of outputting the channel signal, the problem of spectrum mutation at the switching point caused by frequent switching between channels may also occur, resulting in noise at the spectrum mutation point and affecting the listening experience.

[0106] In an improved embodiment, in order to reduce the frequent switching between channels, the method of the application further comprises:

[0107] The channel number selected in the current frame is compared with the channel number output in the last frame, and if they are equal, the channel signal of the channel number selected in the current frame is directly output;

[0108] If not equal, continue outputting the channel signal of the channel number of the last frame output until a preset holding frame number is reached.

[0109] The improved embodiment can realize channel holding and reduce frequent switching between channels. The program implementation example is as follows:

[0110] The current frame, assuming that the channel energy of the channel number index is minimum, is compared with prev_index (representing the last frame energy minimum channel number), if equal, directly output the channel signal of the channel number index. If not equal, reduce 1 from the preset holding frame number channel_swtich_hold (representing the frame number of holding several frames without switching), and continue outputting the channel signal of the channel number prev_index, until channel_swtich_hold is equal to 0, and the channel switching is performed, and the switching flag bit isSwitch is set to 1, that is, the switching flag bit isSwitch is set to true, of course, other flag bit modes can also be used, such as true / false. The channel_swtich_hold is set to a positive initial value (for example, the value 5), and is reset after channel switching and when the channel number index is equal to prev_index.

[0111] In another improved embodiment, in order to avoid the problem of spectral mutation at the switching point after channel switching, the method of the application further comprises:

[0112] After channel switching, the channel signal of the last frame channel number and the channel signal of the current frame channel number are weighted and smoothed frame by frame to obtain a weighted and smoothed speech signal, wherein the weight of the channel signal of the current frame channel number increases frame by frame;

[0113] The weighted and smoothed speech signal is output frame by frame until a preset smoothing frame number is reached.

[0114] The improved embodiment described above can be implemented by the following formula:

[0115] S k,out = (1-w[i])*S k,in (prev_index)+w[i]*S k,in (index) Formula 5

[0116] In formula 5, S k,in (prev_index) represents the Kth frame input speech signal of the last frame channel number, S k,in (index) represents the Kth frame input speech signal of the current frame channel number, and S k,outThe weighted smoothed speech signal of the Kth frame output is represented by w[i], which represents the weight value, and i represents the weight sequence number, i is in the range of 0 to the preset smoothing frame number, and w[i] increases with the increase of i; the preset smoothing frame number is reduced by one each time a frame of the weighted smoothed speech signal is output, until the preset smoothing frame number is equal to 0, and the smoothing switching process is completed;

[0117] The switching flag bit isSwitch is reset to 0, that is, the switching flag bit isSwitch is set to false, and the previous frame channel number prev_index is updated to the current frame channel number index.

[0118] After channel switching, the output signal is obtained by weighted smoothing of the prev_index and index two channel signals using the above formula 5, wherein the weight w[i] is set according to the experiment, such as [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0], and w[i] increases with the increase of i. The weight of the channel signal of the current frame channel number index can be increased frame by frame, the smooth switching between channels is realized, and the spectral mutation at the switching point is avoided.

[0119] Figure 8 The spectral local contrast diagrams of the output signal after channel switching of one embodiment of the application before and after the improvement of channel maintenance and smooth switching are shown, wherein Figure 8 (a) corresponds to before the improvement, Figure 8 (b) corresponds to after the improvement. By comparing Figure 8 (a) and Figure 8 (b), it can be seen that in Figure 8 (a) the position indicated by the arrow, there is a vertical line in the spectrum diagram, indicating that there is spectral mutation at these positions, and in Figure 8 (b) the corresponding position, no vertical line is found in the spectrum diagram, indicating that after the improvement of channel maintenance and smooth switching, the spectral mutation at these positions is eliminated.

[0120] The processing method of the multi-channel wind noise containing speech signal belongs to the same technical concept as the foregoing processing method of the multi-channel wind noise containing speech signal, and the application further provides a processing system of a multi-channel wind noise containing speech signal. Figure 9 The structural schematic diagram of the processing system of the multi-channel wind noise containing speech signal of one embodiment of the application is shown, referring to Figure 9 The system of the application includes:

[0121] The preprocessing module 910 is configured to perform frame division and time-frequency transformation preprocessing on the multi-channel signals input by the plurality of microphones in real time.

[0122] The position grouping module 920 is configured to group the multi-channel signals according to microphone positions, and each channel group comprises at least two channel signals.

[0123] The wind noise detection module 930 is configured to perform wind noise detection on each channel group to determine whether the wind noise is contained in the channel group.

[0124] The beam fusion module 940 is configured to perform beam fusion on all channel signals in the channel group without wind noise frame by frame and output the channel signals after the beam fusion.

[0125] The channel selection module 950 is configured to select the channel signal with the minimum energy or the highest spectral barycenter from all channel signals frame by frame and output the channel signal when it is determined that all channel groups contain wind noise.

[0126] In a preferred embodiment, referring to Figure 10 , Figure 10 FIG. 6 shows a structural schematic diagram of the channel selection module in an embodiment of the present application. The channel selection module 950 comprises:

[0127] The energy screening unit 9501 is configured to select two channels with the minimum energy and the second minimum energy from all channel signals frame by frame when it is determined that all channel groups contain wind noise.

[0128] The energy comparison unit 9502 is configured to divide the energy of the two channels selected by the energy screening unit 9501 to determine whether the quotient is in a preset range close to 1.

[0129] The spectral barycenter factor calculation unit 9503 is configured to calculate the cumulative spectral average barycenter factor of each of the two channels.

[0130] The first channel selection unit 9504 is configured to select the channel signal corresponding to the cumulative spectral average barycenter factor with the larger value to output when the energy comparison unit 9502 determines that the quotient of the energy of the two channels is in the preset range close to 1.

[0131] The second channel selection unit 9505 is configured to select the channel signal with the minimum energy to output when the energy comparison unit 9502 determines that the quotient of the energy of the two channels is not in the preset range close to 1.

[0132] In the above units, the calculation of the spectral barycenter factor calculation unit 9503 is based on the determination of the energy comparison unit 9502 that the quotient of the energy of the two channels is in the preset range close to 1. The first channel selection unit 9504 and the second channel selection unit 9505 are in a mutually exclusive relationship.

[0133] In a preferred embodiment, the spectrum center factor calculation unit 9503 calculates the cumulative spectrum average center factor by the following process:

[0134] The single-frame spectrum center factor is obtained according to the ratio of the frame amplitude frequency product to the frame amplitude value, wherein the frame amplitude frequency product is obtained by summing the products of the amplitudes of all frequency points in each frame and the corresponding frequencies, and the frame amplitude value is obtained by summing the amplitudes of all frequency points in each frame.

[0135] The cumulative spectrum average center factor is obtained by averaging the sum of all single-frame spectrum center factors in the current cumulative frame.

[0136] In an improved embodiment, in order to reduce frequent switching between channels, the system of the present application further comprises:

[0137] The channel maintaining module is configured to compare the channel number selected in the current frame with the channel number output in the last frame during the output of the channel signal, and if they are equal, directly output the channel signal of the channel number selected in the current frame, and if they are not equal, continue to output the channel signal of the channel number output in the last frame until a preset number of maintained frames is reached.

[0138] In another improved embodiment, in order to avoid the problem of spectrum mutation at the switching point after channel switching, the system of the present application further comprises:

[0139] The smooth switching module is configured to, after channel switching, perform weighted smoothing on the channel signal of the channel number of the last frame and the channel signal of the channel number of the current frame frame by frame to obtain a weighted smoothed speech signal, wherein the weight of the channel signal of the channel number of the current frame increases frame by frame, and output the weighted smoothed speech signal frame by frame until a preset number of smoothed frames is reached.

[0140] The implementation process of each module in the multi-channel wind noise containing speech signal processing system of the present application can be referred to the method embodiments described above, which will not be described here.

[0141] The multi-channel wind noise containing speech signal processing method and system described above belong to the same technical concept, and the present application further provides an electronic device. Figure 11 An electronic device provided for an embodiment of the present application is shown in the structural diagram. Referring to Figure 11 The electronic device provided by the present application comprises a plurality of microphones, a memory and a processor, wherein the plurality of microphones respectively collect speech signals of the surrounding environment in real time to obtain multi-channel signals; the memory stores a computer program, the computer program is loaded and executed by the processor to implement the multi-channel wind noise containing speech signal processing method described above, which will not be described here.

[0142] At the hardware level, the electronic device can also selectively include a display panel, an interface module, a communication module, a speaker, and the like. The memory, the processor, and the display panel, the interface module, the communication module, the speaker, the plurality of microphones, and the like can be connected to each other through an internal bus. The internal bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, and the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 10 Only one bidirectional arrow is used to represent the bus in the middle, but it does not mean that there is only one bus or only one type of bus.

[0143] The application further provides a computer readable storage medium storing one or more computer programs, which, when executed by a processor, implement the aforementioned multi-channel wind noise containing speech signal processing method, which will not be described again here.

[0144] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, an electronic device, or a computer program product. Therefore, the present application can be in the form of a completely hardware embodiment, a completely software embodiment, or a combination of software and hardware. Moreover, the present application can be in the form of a computer program product implemented on one or more computer readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing a computer program.

[0145] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles, or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or further include elements inherent in such processes, methods, articles, or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article, or device including the element.

[0146] The above is only an embodiment of the present application and is not intended to limit the present application. The present application can have various changes and modifications for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application shall be included in the scope of the claims of the present application.

Claims

1. A method of processing a multi-channel speech signal containing wind noise, characterized by, The method comprises: frame dividing and time-frequency transformation preprocessing are performed on multiple microphone real-time input multi-channel signals respectively; the multi-channel signals are grouped according to microphone positions, and each channel group comprises at least two channel signals; wind noise detection is performed on each channel group in units of channel groups to determine whether wind noise is contained; when it is determined that all channel groups or only part of the channel groups do not contain wind noise, all channel signals in the channel group without wind noise are output after frame-by-frame beam fusion; when it is determined that all channel groups contain wind noise, the channel signal with the minimum energy or the highest spectral barycenter is selected from all channel signals frame by frame according to the time domain energy size and the frequency domain spectral distribution of the channel; the channel signal with the minimum energy or the highest spectral barycenter is selected from all channel signals frame by frame according to the time domain energy size and the frequency domain spectral distribution of the channel, comprising: two channels with the minimum energy and the second minimum energy are selected from all channel signals frame by frame, the energy of the two channels is divided, if the quotient is in a preset range close to 1, the cumulative spectral average barycenter factor of each of the two channels is calculated, and the channel signal corresponding to the cumulative spectral average barycenter factor with the larger value is selected for output, otherwise, the channel signal with the minimum energy is selected for output.

2. The method of claim 1, wherein, The method further comprises: during the output of the channel signal, the channel number selected in the current frame is compared with the channel number output in the last frame, if they are equal, the channel signal of the channel number selected in the current frame is directly output; if they are not equal, the channel signal of the channel number output in the last frame is continuously output until a preset number of frames is reached.

3. The method of claim 2, wherein, The method further comprises: after channel switching, the channel signal of the channel number in the last frame is frame by frame weighted and smoothed with the channel signal of the channel number in the current frame, to obtain a weighted and smoothed speech signal, wherein the weight of the channel signal of the channel number in the current frame increases frame by frame; the weighted and smoothed speech signal is output frame by frame until a preset number of smoothing frames is reached.

4. The method of claim 1, wherein, The calculation process of the cumulative spectral average barycenter factor comprises: a single-frame spectral barycenter factor is obtained according to the ratio of the frame amplitude frequency product to the frame amplitude value, wherein the frame amplitude frequency product is obtained by summing the product of the amplitude value of each frequency point and the corresponding frequency in each frame, and the frame amplitude value is obtained by summing the amplitude value of each frequency point in each frame; the sum of all single-frame spectral barycenter factors in the current cumulative frame is averaged to obtain the cumulative spectral average barycenter factor.

5. A system for processing a multi-channel wind-noisy speech signal, characterized by The system comprises: a preprocessing module configured to perform frame dividing and time-frequency transformation preprocessing on multiple microphone real-time input multi-channel signals respectively; a position grouping module configured to group the multi-channel signals according to microphone positions, and each channel group comprises at least two channel signals; a wind noise detection module configured to perform wind noise detection on each channel group in units of channel groups to determine whether wind noise is contained; a beam fusion module configured to output all channel signals in the channel group without wind noise after frame-by-frame beam fusion when it is determined that all channel groups or only part of the channel groups do not contain wind noise; The channel selection module is configured to select, from all channel signals, a channel signal with minimum energy or a channel signal with the highest spectral barycenter when it is determined that all channel groups contain wind noise, and output the channel signal frame by frame. The channel selection module comprises: The energy screening unit is configured to select, from all channel signals, two channels with minimum energy and sub-minimum energy when it is determined that all channel groups contain wind noise. The energy comparison unit is configured to divide the energy of the two channels selected by the energy screening unit, and determine whether the quotient is within a preset range close to 1. The spectral barycenter factor calculation unit is configured to calculate the cumulative spectral average barycenter factor of each of the two channels. The first channel selection unit is configured to select, when the energy comparison unit determines that the quotient of the energy of the two channels is within the preset range close to 1, the channel signal corresponding to the cumulative spectral average barycenter factor with the larger value to output. The second channel selection unit is configured to select, when the energy comparison unit determines that the quotient of the energy of the two channels is not within the preset range close to 1, the channel signal with minimum energy to output.

6. The system of claim 5, wherein, The system further comprises: The channel retention module is configured to compare the channel number selected in the current frame with the channel number output in the last frame during output of the channel signal, and if they are equal, directly output the channel signal of the channel number selected in the current frame; if they are not equal, continue to output the channel signal of the channel number output in the last frame until a preset number of frames are reached.

7. The system of claim 6, wherein, The system further comprises: The smooth switching module is configured to, after channel switching, perform weighted smoothing on the channel signal of the channel number in the last frame and the channel signal of the channel number in the current frame frame by frame to obtain a weighted smoothed speech signal, wherein the weight of the channel signal of the channel number in the current frame increases frame by frame; and output the weighted smoothed speech signal frame by frame until a preset number of smoothing frames are reached. 8.An electronic device comprising a plurality of microphones, a memory, and a processor, The plurality of microphones are configured to collect speech signals of a surrounding environment in real time to obtain multi-channel signals. The memory stores a computer program, and the computer program is loaded and executed by the processor to implement the processing method of the multi-channel wind noise containing speech signal according to any one of claims 1-4. 9.A computer readable storage medium storing one or more computer programs, which, when executed by a processor, implement the processing method of the multi-channel wind noise containing speech signal according to any one of claims 1-4.

Citation Information

Patent Citations

  • Wind noise suppression method and system suitable for artificial cochlea

    CN111261182A

  • Noise reduction in an audio system

    US10192566B1