Distributed microphone array noise reduction method, device and control equipment

By receiving the signal strength information of the smart device and calculating the time delay, the microphone array structure of the target device in the distributed microphone array is determined, which solves the problem of voice noise reduction in the distributed microphone array and realizes efficient voice signal processing.

CN114613382BActive Publication Date: 2025-05-06BEIJING INTENGINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210092715.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-26
Publication Date
2025-05-06
Estimated Expiration
2042-01-26

AI Technical Summary

Technical Problem

In a distributed microphone array scenario, it is difficult to determine the microphone position that receives user voice commands, thereby achieving effective voice noise reduction.

Method used

By receiving the signal strength information of each intelligent device, the target device to be interacted with by the user is determined, and the microphone array structure of the target device is calculated according to the time delay between the target device and the auxiliary device, thereby performing noise reduction operations.

Benefits of technology

It realizes accurate identification of target devices and determination of microphone array structure in distributed microphone arrays, and can accurately perform voice noise reduction regardless of the device position change, solving the application problem of voice noise reduction algorithm in distributed arrays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114613382B_ABST
    Figure CN114613382B_ABST
Patent Text Reader

Abstract

The present invention discloses a distributed microphone array noise reduction method, device and control device, the method is applied to the control device in the target area, the target area also includes a plurality of smart devices with microphone arrays, and the smart devices include at least one auxiliary device with a known microphone array structure; the method includes: receiving the signal strength information reported by each smart device for the user's voice command, and based on each signal strength information, determining the target device to be interacted with by the user from each smart device; identifying the time delay between the target device and the auxiliary device for receiving the user's voice command, and calculating the microphone array structure of the target device according to the time delay and the microphone array structure of the auxiliary device; performing noise reduction operation on the user's voice command based on the microphone array structure of the target device. The technical solution provided by the present invention realizes the function of determining the position of the microphone receiving the user's command in the distributed microphone array.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of signal processing, and in particular to a distributed microphone array noise reduction method, device and control equipment. Background Art

[0002] With the popularity of smart homes, more and more families have more and more smart devices with voice interaction functions, such as smart TVs and smart speakers placed in the living room, smart table lamps placed in the bedroom, smart refrigerators placed in the kitchen, etc. These devices are equipped with a single microphone or microphone array to realize the voice interaction function, and these smart devices are distributed. Unlike regular microphone arrays, multiple microphones in a distributed microphone array belong to different devices, and the number and position of devices are not fixed, that is, the geometric shape and spatial position of the distributed microphone array are unpredictable, and the size of the device is usually not large and can be moved. Therefore, beamforming and other noise reduction algorithms that perform conventional array signal processing based on the position of the microphone array are difficult to implement. Therefore, how to determine the position of the microphone that receives the user's command in the distributed microphone array, so as to reduce the noise of the user's voice command, is a problem that must be solved. Summary of the invention

[0003] In view of this, the embodiments of the present invention provide a distributed microphone array noise reduction method, apparatus and control device, thereby realizing the function of determining the position of the microphone receiving the user command in the distributed microphone array.

[0004] According to the first aspect, the present invention provides a distributed microphone array noise reduction method, which is applied to a control device in a target area, wherein the target area also includes multiple smart devices with microphone arrays, and the smart devices include at least one auxiliary device with a known microphone array structure; the method includes: receiving signal strength information reported by each of the smart devices in response to a user voice command, and based on each of the signal strength information, determining a target device with which the user is to interact from each of the smart devices; identifying a time delay between the target device and the auxiliary device in receiving the user voice command, and calculating the microphone array structure of the target device based on the time delay and the microphone array structure of the auxiliary device; and performing noise reduction operations on the user voice command based on the microphone array structure of the target device.

[0005] Optionally, based on each of the signal strength information, determining the target device with which the user wants to interact from each of the smart devices includes: comparing the signal strength values ​​represented by each of the signal strength information, and obtaining the target device information corresponding to the maximum signal strength value; based on the target device information, sending a wake-up command to the corresponding target device, so that the target device can interact with the user.

[0006] Optionally, the noise reduction operation on the user voice command based on the microphone array structure of the target device includes: calculating the sound source position of the user voice command based on the microphone array structure of the target device; time-aligning the user voice commands received by each microphone in the microphone array of the target device based on the microphone array structure of the target device and the sound source position; and weightedly summing the aligned user voice commands to obtain a noise-reduced target user voice command.

[0007] Optionally, the method of calculating the sound source position of the user voice command based on the microphone array structure of the target device includes: pairing the microphones in the microphone array of the target device in pairs to obtain multiple microphone pairing combinations; dividing the azimuth angle range and elevation angle range to be observed into grids with preset intervals, and combining them in pairs to obtain multiple angle observation grids; traversing and calculating the angle spectrum function corresponding to each microphone pairing combination based on the user voice command; obtaining the angle observation grid corresponding to the target angle spectrum function value in space to obtain the sound source position of the user voice command, the target angle spectrum function value being the maximum function value in the angle spectrum function corresponding to each microphone pairing combination.

[0008] Optionally, the method of calculating the sound source position of the user voice command based on the microphone array structure of the target device also includes: dividing the user voice command into a plurality of equally spaced signal blocks, each signal block being divided into a plurality of equally spaced signal frames; traversing each signal block to generate a plurality of candidate angle observation grids by the steps of traversing and calculating the angle spectrum function corresponding to each microphone pairing combination based on the user voice command to obtaining the angle observation grid corresponding to the target angle spectrum function value in space; determining a second angle observation grid from the plurality of candidate angle observation grids based on data smoothing processing to obtain the sound source position of the user voice command; wherein the target angle spectrum function value of each signal block is generated based on the maximum value of the angle spectrum function of the signal frame within the signal block.

[0009] Optionally, the noise reduction operation on the user voice command based on the microphone array structure of the target device also includes: inputting each user voice command after time alignment processing into a blocking matrix, and outputting the corresponding noise signal in each user voice command; inputting the target user voice command and multiple noise signals into a preset adaptive filter at the same time, so as to use multiple noise signals to perform noise reduction operation on the user voice command after weighted summation processing, and obtain a second target user voice command.

[0010] Optionally, the method further includes: performing time synchronization between the control device and each smart device.

[0011] According to the second aspect, the present invention provides a distributed microphone array noise reduction device, which is applied to a control device in a target area, wherein the target area also includes multiple smart devices with microphone arrays, and the smart devices include at least one auxiliary device with a known microphone array structure. The device includes: an interactive device determination module, which receives signal strength information reported by each of the smart devices in response to a user voice command, and determines a target device with which the user is to interact from each of the smart devices based on each of the signal strength information; a microphone array estimation module, which identifies the time delay between the target device and the auxiliary device in receiving the user voice command, and calculates the microphone array structure of the target device based on the time delay and the microphone array structure of the auxiliary device; and a signal noise reduction module, which performs noise reduction operations on the user voice command based on the microphone array structure of the target device.

[0012] According to the third aspect, an embodiment of the present invention provides a control device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the method described in the first aspect or any optional implementation manner of the first aspect by executing the computer instructions.

[0013] According to a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method described in the first aspect or any optional implementation manner of the first aspect.

[0014] The technical solution provided by this application has the following advantages:

[0015] The technical solution provided by the present application is to set a control device and an auxiliary device of a known microphone array structure in a target area containing multiple smart devices. When the user issues a voice command, the signal strength information is calculated according to the voice command received by each smart device, and then each smart device sends the signal strength information to the control device. The control information finds the target device corresponding to the voice command with the strongest signal strength from the smart device according to the signal strength information, that is, accurately identifies the target device closest to the user. Then, the time delay between the target device and the auxiliary device for receiving the user's voice command is identified. According to the time delay and the known microphone array structure of the auxiliary device, a coordinate equation can be established to solve the spatial coordinates of each microphone device in the microphone array of the target device, and obtain the microphone array structure of the target device. Through the above steps, no matter how the position of the smart device changes, the microphone array structure of the smart device can be accurately obtained, and then the user's voice command can be denoised based on the microphone array structure of the target device combined with microphone noise reduction algorithms such as beamforming. Thereby solving the problem that the voice noise reduction algorithm is difficult to apply in the distributed microphone array scenario.

[0016] In addition, when the embodiment of the present invention performs noise reduction on the voice signal, the specific coordinate position of the sound source in space is first determined based on the microphone array structure of the target device, and then the user voice command signals with different time delays received by multiple microphones in the microphone array can be time-aligned according to the position of the sound source, so as to weight the sum of the user command voice signals received by each microphone. In the weighted summation process, based on the random characteristics of white noise, the noise will not be amplified, so that the noise ratio in the enhanced user voice command signal is greatly reduced, achieving the purpose of fast and simple noise reduction. In addition, based on the blocking matrix, multiple noise signals are extracted, and the enhanced user voice command is denoised using a filter, further improving the signal quality of the user voice command. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present invention in any way. In the accompanying drawings:

[0018] Figure 1 A schematic diagram of the steps of a distributed microphone array noise reduction method in one embodiment of the present invention is shown;

[0019] Figure 2 A schematic diagram of an Ad-hoc network structure in one embodiment of the present invention is shown;

[0020] Figure 3 An example diagram of the spatial structure of a smart device and an auxiliary array in one embodiment of the present invention is shown;

[0021] Figure 4A schematic diagram showing the relationship between the time delay of a user voice command received by a target device and an auxiliary array in one embodiment of the present invention is shown;

[0022] Figure 5 A schematic structural diagram of a distributed microphone array noise reduction device in one embodiment of the present invention is shown;

[0023] Figure 6 A schematic structural diagram of a control device in one embodiment of the present invention is shown. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0025] See also Figure 1 In one embodiment, a distributed microphone array noise reduction method is applied to a control device in a target area, wherein the target area also includes a plurality of smart devices with microphone arrays, and the smart devices include at least one auxiliary device with a known microphone array structure, and specifically includes the following steps:

[0026] Step S101: receiving signal strength information reported by each smart device in response to a user voice command, and determining a target device with which the user is to interact from each smart device based on each signal strength information.

[0027] Step S102: Identify the time delay between the target device and the auxiliary device in receiving the user voice command, and calculate the microphone array structure of the target device according to the time delay and the microphone array structure of the auxiliary device.

[0028] Step S103: performing noise reduction operation on the user voice command based on the microphone array structure of the target device.

[0029] Specifically, the prerequisite for array signal processing algorithms based on beamforming and other technologies is that the microphone array configuration is known, so that the corresponding steering vector can be generated according to the array configuration for subsequent algorithms. In a distributed microphone array, the configuration of the microphone array of each device is unknown, so the array signal processing algorithm cannot be directly applied. In an embodiment of the present invention, all smart devices in the target area are equipped with microphone arrays, and the microphone array of each smart device includes an audio acquisition module, a wireless transmission module and a network device clock module. Its microphone array can be a regular array: such as a linear array, a planar array, a circular array, a stereo array, etc., or an irregular array; the microphone array of the auxiliary device can be a regular array or an irregular array; if the structural information of the microphone array in any smart device is known, that is, the array configuration and the microphone spacing are known, then the smart device can be used as an auxiliary device, and no additional auxiliary device needs to be configured. If the structural information of the microphone array of each device is unknown, it is also necessary to configure an auxiliary device with known microphone array structural information; in this embodiment, the control device uses the received information to perform microphone array position estimation, sound source localization, speech noise reduction, speech recognition and other processing. The communication connection method includes but is not limited to wireless communication methods such as wifi, bluetooth, zigbee, etc., and also includes wired communication methods such as network cables. In this embodiment, Figure 2 As shown, an Ad-hoc network is established between smart devices and control devices to achieve communication connection between devices, and also solve the problem of command interference between smart devices.

[0030] In a specific embodiment, if Figure 3The figure shows a lighting control scene in a room (i.e., the target area), which includes three smart devices with the same wake-up word, namely, a ceiling lamp, a floor lamp, and a table lamp, each of which has its own microphone array. If the device is not connected to the Ad-hoc network, the three devices will be woken up at the same time when the user issues a wake-up command, which greatly affects the user experience; and when these devices are connected to the Ad-hoc network, the control device receives the signal strength information sent by each device, and the signal strength information can be used to characterize the distance between the user and each device. The closer the user is to a certain device, the stronger the signal strength received by the device should be. Signal strength is usually measured by signal power and energy. In this embodiment, the short-time energy (i.e., the amplitude parameter of the signal in a short time) sent by each device to characterize the signal strength is compared. The closer the distance to the user, the smaller the propagation attenuation, and the larger the amplitude of the sound signal received by the microphone, i.e., the larger the short-time energy; on the contrary, the farther the distance to the user, the greater the propagation attenuation, and the smaller the amplitude of the sound signal received by the microphone, i.e., the smaller the short-time energy. By finding the device corresponding to the signal with the largest short-time energy, the table lamp closest to the user can be used as the target device. The user only needs to interact with it to control the floor lamp farthest away. Therefore, the problem of interaction confusion caused by multiple devices receiving user voice commands at the same time is also solved. At the same time, the positions of these three devices may not be fixed. For example, floor lamps and table lamps are easy to move, which means that the accurate spatial position information of their microphone arrays cannot be obtained by conventional measurement methods, so the commonly used array signal processing algorithms cannot be used to perform noise reduction operations on user commands. Since the structural information of the microphone arrays of these three devices is unknown, an auxiliary array of three microphone uniform linear arrays with known array element spacing is configured in the corner of the room ceiling as an auxiliary device. The auxiliary array can be used to obtain the spatial position information of the microphone array of the wake-up device, so that the conventional array signal processing algorithm can be used for noise reduction. If the microphone array structure of at least one of the three smart devices is known, the device with the known microphone array structure can be used as an auxiliary device, and no additional auxiliary device needs to be configured. In addition, determining the spatial position of the device microphone array is essentially to calculate the three-dimensional coordinates of the microphone array. Therefore, at least three microphones are required to form an auxiliary array of the auxiliary device to obtain three equations to solve three unknowns. In this embodiment, the array form of three microphones uniform linear array, which is the easiest to calculate, is adopted. The auxiliary arrays of other array forms can also achieve this purpose, but the array form and spatial position calculation are more complicated. Finally, after obtaining the microphone array structure of the target device, the coordinates of each microphone in the microphone array in space and the positional relationship between the microphones are obtained, so that the commonly used beamforming noise reduction algorithm can be used to perform noise reduction operations on the received user voice commands.

[0031] In a specific embodiment, after the control device sends the wake-up command, the auxiliary array receiving signal is uploaded to the control device through the communication network. At this time, the control device has both the auxiliary array receiving signal and the target device receiving signal. The relative delay relationship between the received auxiliary array and the target device user voice command signal is as follows: Figure 4 As shown, the time alignment of the signals received by all microphones is performed based on the microphone 1 of the auxiliary array. Figure 4 It can be seen that there is a large time delay between the target device and the auxiliary array, which is determined by the distance difference of the sound signal propagating in space. Based on this distance difference and the microphone structure inside the auxiliary array, the structure of the microphone array inside the target device can be estimated, that is, the three-dimensional coordinates of each microphone of the target device can be obtained. When the auxiliary array is a uniform linear array, there are 3 microphones with a microphone spacing of d. The position of microphone 1 is recorded as the origin of the coordinate system, that is, the three-dimensional coordinates of microphone 1 are (0,0,0), the three-dimensional coordinates of microphone 2 are (d,0,0), and the three-dimensional coordinates of microphone 3 are (2d,0,0). If the target device has 4 microphones, the following formula can be established based on the time delay between the target device and the auxiliary array:

[0032]

[0033] Among them, x MICn ,y MICn and z MICn Respectively represent the horizontal coordinate, vertical coordinate and vertical coordinate of the nth (n=1, 2, 3, 4) microphone of the target device, τ n1 , τ n2 and τ n3 They respectively represent the time delay between the nth microphone of the wake-up device and the auxiliary array microphone 1, microphone 2 and microphone 3. By solving the above expressions, the three-dimensional coordinates of each microphone in the target device, that is, the microphone array structure of the target device, can be obtained.

[0034] Specifically, in one embodiment, the above step S101 specifically includes the following steps:

[0035] Step 1: Compare the signal strength values ​​represented by each signal strength information, and obtain the target device information corresponding to the maximum signal strength value.

[0036] Step 2: Send a wake-up instruction to the corresponding target device based on the target device information, so that the target device can interact with the user.

[0037] Specifically, in this embodiment, based on the short-time energy in the above steps S101 to S102 as a measure of the voice signal strength, the formula for calculating the short-time energy of the user's voice command by any smart device is as follows:

[0038]

[0039] Where P represents the short-time energy of the device receiving the user's voice command, N represents the number of microphones of the device, M represents the number of sampling points of the received signal voice segment, and x n (m) refers to the amplitude of the mth sampling point of the nth microphone.

[0040] Compare the size of the short-time energy sent by each smart device, and then control the device to obtain the target device information with the largest short-time energy, including but not limited to the device ID and MAC code, and then send the wake-up command to the target device corresponding to the target device information, so that the target device can continue to interact with the user to control any device in the target area.

[0041] Specifically, in one embodiment, the above step S103 specifically includes the following steps:

[0042] Step 3: Calculate the sound source position of the user's voice command based on the microphone array structure of the target device.

[0043] Step 4: Based on the microphone array structure and sound source position of the target device, time-align the user voice commands received by each microphone in the microphone array of the target device.

[0044] Step 5: Perform weighted summation on the aligned user voice commands to obtain the noise-reduced target user voice command.

[0045] Specifically, this embodiment first determines the direction angle and distance of the sound source according to the geometric relationship between the sound source and the microphone array of the target device, and then designs the fractional-order FIR filter corresponding to each microphone of the target device based on the time difference from the sound source to each array element in the microphone array. In order to improve the signal processing effect, in order to approximate the non-stationary speech signal to a stationary signal, this embodiment also performs frame processing on the user voice command. After that, for each frame signal, with any microphone as the reference, such as microphone 1 in the target device, the signal is delayed according to the fractional-order FIR filter corresponding to each microphone, that is, the time alignment of the remaining microphone receiving signals with the microphone 1 receiving signal is completed. The multiple delayed signals obtained are weighted and summed (i.e., beamforming) to enhance the useful signal, while the white noise will not be enhanced due to the random characteristics of weighting, thereby achieving preliminary suppression of spatial noise and obtaining the target user voice command with noise reduction. The formula for weighted summation of multiple delayed signals is as follows:

[0046]

[0047] Among them, τ nrepresents the delay of waking up the nth microphone of the device relative to microphone 1, FBF represents fixed beam forming, x FBF represents the fixed beamforming signal, x n Represents the received signal of the nth microphone. The subtraction operation defines the sampling point m to be shifted in the positive direction of the time axis.

[0048] Specifically, in one embodiment, the above step three specifically includes the following steps:

[0049] Step 6: Pair the microphones in the microphone array of the target device in pairs to obtain multiple microphone pairing combinations.

[0050] Step 7: Divide the azimuth angle range and elevation angle range to be observed into grids with preset intervals, and combine them in pairs to obtain multiple angle observation grids.

[0051] Step 8: Based on the user voice command, traverse and calculate the angle spectrum function corresponding to each microphone pairing combination.

[0052] Step nine: Obtain the angle observation grid corresponding to the target angle spectrum function value in space to obtain the sound source position of the user's voice command. The target angle spectrum function value is the maximum function value in the angle spectrum function corresponding to each microphone pairing combination.

[0053] Specifically, after obtaining the microphone array structure inside the target device, the direction of arrival is estimated based on the obtained microphone array structure of the target device and the voice signal, that is, the user is positioned using the relative time delay between the microphones inside the target device, that is, the user's azimuth and elevation angle information is obtained.

[0054] In this embodiment, the azimuth and elevation information of the user is obtained by calculating the angle spectrum function. Assuming that the target device includes 4 microphones, first, pair the microphones of the target device. Since the wake-up device has a total of 4 microphones, the microphones are paired in pairs, and there are 6 non-repetitive microphone combinations, namely microphone 1 and microphone 2, microphone 1 and microphone 3, microphone 1 and microphone 4, microphone 2 and microphone 3, microphone 2 and microphone 4, and microphone 3 and microphone 4; secondly, the azimuth angle range and pitch angle range to be observed are grid-divided in space, such as the azimuth angle is divided into 5° intervals within the range of 0° to 90°, and a total of 19 angle values ​​(including 0° and 90°) are obtained. The pitch angle is 0°. The angle is divided into 5° intervals within 90° to obtain a set of 19 angle values. The azimuth and pitch angles are combined in pairs, and a total of 361 angle combinations are obtained. The estimated value of the microphone array structure of the target device is used to obtain the delay set τ corresponding to each pair of microphone combinations. The user voice command signal is subjected to short-time Fourier transform to convert the time domain signal into the frequency domain. Since the voice frequency during speaking is generally concentrated in the range of 300Hz to 3400Hz, only the data in this frequency band is selected for subsequent calculations to further improve the accuracy of sound source localization.

[0055] The steps for calculating the angle spectrum function based on the time delay set τ are as follows:

[0056] (1) Define the normalized diffuse noise covariance matrix:

[0057]

[0058] Among them, f represents the frequency point (that is, the frequency value corresponding to each value after the signal is transformed into the frequency domain by short-time Fourier transform), represents the distance between microphone n1 and microphone n2, c represents the speed of sound, and sinc(·) represents the Sinker function.

[0059] (2) Combine the normalized diffuse noise covariance matrix in step (1) to define the noise projection matrix Α and the signal / noise subspace power concatenation matrix Λ:

[0060] The column vectors of the matrix Α are represented by a(f,τ)a H (f,τ)ψ -1 (f), a(f,τ) represents the frequency-domain steering vector obtained from the delay set, (·) H It means to find the conjugate transposed matrix of a matrix, (·) -1 It means to find the inverse matrix of the matrix, and the matrix Λ is:

[0061] Λ=[diag(A H a(f,τ)a H (f,τ)A),diag(AH ψ(f)A)]

[0062] Where diag(·) means finding the diagonal elements of the matrix.

[0063] (3) Calculate the neighborhood empirical covariance matrix at the time-frequency point (t, f) (i.e., the frequency point f of the signal at the current time t):

[0064]

[0065] Where win(·) represents the window function, X(t′,f′) represents the data at the time-frequency point (t′,f′);

[0066] (4) Based on the matrix of step (2) and step (3), the signal energy v of the user's voice command is calculated by substituting the matrix into the formula: s (t,f,τ) and noise signal energy v b (t,f,τ):

[0067] [v s (t,f,τ),v b (t,f,τ)] T =Λ -1 diag(A H R(t,f)A)

[0068] (5) Therefore, the time-frequency angle spectrum function of the current microphone pairing combination can be expressed as:

[0069]

[0070] In the selected frequency band, not all frequency points can obtain correct angle estimation results. Therefore, in order to further reduce the impact of noise on the direction finding result, this embodiment adopts a frequency power weighted method to comprehensively obtain the time domain angle spectrum function:

[0071]

[0072] Among them, P f Indicates the power value at frequency f.

[0073] Repeat steps (1) to (5) for all microphone pairing combinations, accumulate the time domain angle spectrum function and find the maximum value to obtain the final angle spectrum function:

[0074]

[0075] Finally, the angle observation grid corresponding to the maximum value of the angle spectrum function in space is found, and the corresponding azimuth and elevation angle combination can be obtained as the estimated value of the user's sound source position.

[0076] Specifically, in one embodiment, the above-mentioned step S103 further includes the following steps:

[0077] Step Ten: Divide the user voice command into multiple equally spaced signal blocks, and each signal block is divided into multiple equally spaced signal frames.

[0078] Step Eleven: Based on Steps Eight to Nine, traverse each signal block to generate multiple candidate angle observation grids;

[0079] Step Twelve: Determine the second angle observation grid from multiple candidate angle observation grids based on data smoothing processing to obtain the sound source position of the user voice command. Among them, the target angle spectrum function value of each signal block is generated based on the maximum value of the angle spectrum function of the signal frames within the signal block.

[0080] Specifically, in order to further improve the accuracy of user sound source position localization, the control device divides the received signal of the target device into frames with a frame length of L, and splices J (J < L) frames into a data block; then performs a short-time Fourier transform on each frame signal in each data block to convert the time-domain signal to the frequency domain, and then repeats the operations of the above Steps Eight to Nine for each frame signal to obtain the angle spectrum function corresponding to each frame signal, and then calculates the maximum value of the angle spectrum function of each frame signal, and then finds the maximum value (hereinafter referred to as the second maximum value) among the maximum values of the angle spectrum functions calculated by each frame signal, and then uses the second maximum value as the angle spectrum function value of the current signal block. Finally, smooth processing is performed on the estimated values of all signal blocks, including but not limited to mean filtering, median filtering, and Gaussian filtering, to remove the noise and distortion in the estimated values, and the estimated value of the user position can be obtained, that is, the azimuth angle estimate and the elevation angle estimate. In this embodiment, median filtering is used to smooth the estimated values of all signal blocks. For example, the measured estimated value sequence (azimuth angle, elevation angle) is (30, 30) (30, 30) (0, 0) (30, 30) (30, 30), and (0, 0) is a bad value affecting parameter estimation. After median filtering, the sequence becomes (30, 30) (30, 30) (30, 30) (30, 30) (30, 30), thereby improving the direction finding stability. Through the above Steps Ten to Eleven, the signal is frame-processed to improve the smoothness of non-stationary voice signals. Then, the maximum value of the angle spectrum function estimated value in each frame is used as the angle spectrum function value of the entire data block, which is more accurate than the estimation result of the entire segment of the signal, thereby improving the accuracy of sound source localization.

[0081] Specifically, in one embodiment, the above-mentioned step S103 further includes the following steps:

[0082] Step Thirteen: Input the time-aligned user voice commands into a blocking matrix and output the corresponding noise signals in each user voice command.

[0083] Step 14: Input the target user voice command and multiple noise signals into a preset adaptive filter at the same time, so as to use the multiple noise signals to perform a noise reduction operation on the user voice command after weighted summation processing, and obtain a second target user voice command.

[0084] Specifically, in order to further improve the fidelity of user voice commands, this embodiment further performs noise reduction processing on user voice commands. For steps three to four, the microphone array beam is pointed in the direction of the user, and the spatial noise is preliminarily suppressed through beamforming to obtain a preliminarily enhanced voice signal. Each original signal is passed through a blocking matrix to eliminate the desired voice signal in the noisy voice signal, and only the noise signal is output. Finally, the preliminarily enhanced voice signal and noise signal are passed through an adaptive noise canceller (i.e., an adaptive filter) to eliminate the noise signal in the preliminarily enhanced voice signal, obtain a clean voice signal, and complete the noise reduction, so that the user voice command features for subsequent voice recognition and other operations are more prominent and have higher fidelity. The specific operations are as follows:

[0085] First, input each user's voice command after time alignment in step 4 into the blocking matrix to obtain the noise signal. The blocking matrix is

[0086]

[0087] Then, the target user voice command and noise signal are eliminated by an adaptive noise canceller. The adaptive noise canceller is a set of adaptive filters, and its weight coefficient is updated by the Normalized Least Mean Square (NLMS) algorithm. For the kth (k = 1, 2, 3) noise signal, the NLMS algorithm includes two steps: error estimation and filter tap weight coefficient update:

[0088] e k (l) = d(l) - ω k T (l)x k (l)

[0089]

[0090] Among them, l represents the frame number, x k represents the k-th noise signal, ω k represents the weight vector of the kth filter, d represents the initial enhanced speech signal, e k represents the k-th error signal, represents the step size factor, and η represents the minimum value to prevent the denominator from being zero. Finally, the three error signals are summed and then subtracted from the initially enhanced speech signal to obtain the noise reduction signal of the current frame; finally, all frames are traversed to obtain the user voice command after noise reduction, that is, the second target user voice command.

[0091] Specifically, in one embodiment, a distributed microphone array noise reduction method provided by an embodiment of the present invention further includes the following steps:

[0092] Step 15: Synchronize the time between the control device and each smart device.

[0093] Specifically, because the clock source of the microphone array of each smart device is different, but the time when the user speaks is fixed, for example, the user speaks at time a, device 1 should receive the sound signal at time b and device 2 should receive the sound signal at time c. The direction θ based on this time is the direction of the user's sound source. However, since the clock source is not synchronized, the receiving time of device 1 is b, but the receiving time of device 2 may be (c+δc). Therefore, the direction corresponding to this time difference is not necessarily θ. If time synchronization between devices is not performed, inaccurate positioning of the user's sound source may occur. Therefore, this embodiment performs time synchronization of all devices to ensure that the reference time of each smart device is the same, thereby further improving the accuracy of user sound source positioning.

[0094] Through the above steps, the technical solution provided by the present application sets a control device and an auxiliary device of a known microphone array structure in a target area containing multiple smart devices. When the user issues a voice command, the signal strength information is calculated according to the voice command received by each smart device, and then each smart device sends the signal strength information to the control device. The control information finds the target device corresponding to the voice command with the strongest signal strength from the smart device according to the signal strength information, that is, accurately identifies the target device closest to the user. Then identify the time delay between the target device and the auxiliary device for receiving the user's voice command. According to the time delay and the known microphone array structure of the auxiliary device, a coordinate equation can be established to solve the spatial coordinates of each microphone device in the microphone array of the target device, and obtain the microphone array structure of the target device. Through the above steps, no matter how the position of the smart device changes, the microphone array structure of the smart device can be accurately obtained, and then the user's voice command can be denoised based on the microphone array structure of the target device combined with microphone noise reduction algorithms such as beamforming. Thereby solving the problem that the voice noise reduction algorithm is difficult to apply in the distributed microphone array scenario.

[0095] In addition, when the embodiment of the present invention performs noise reduction on the voice signal, the specific coordinate position of the sound source in space is first determined based on the microphone array structure of the target device, and then the user voice command signals with different time delays received by multiple microphones in the microphone array can be time-aligned according to the position of the sound source, so as to weight the sum of the user command voice signals received by each microphone. In the weighted summation process, based on the random characteristics of white noise, the noise will not be amplified, so that the noise ratio in the enhanced user voice command signal is greatly reduced, achieving the purpose of fast and simple noise reduction. In addition, based on the blocking matrix, multiple noise signals are extracted, and the enhanced user voice command is denoised using a filter, further improving the signal quality of the user voice command.

[0096] like Figure 5 As shown, this embodiment also provides a distributed microphone array noise reduction device, which is applied to a control device in a target area. The target area also includes a plurality of smart devices with microphone arrays, and the smart devices include at least one auxiliary device with a known microphone array structure. The device includes:

[0097] The interactive device determination module 101 receives the signal strength information reported by each smart device for the user's voice command, and determines the target device to be interacted with from each smart device based on each signal strength information. For details, please refer to the relevant description of step S101 in the above method embodiment, which will not be repeated here.

[0098] The microphone array estimation module 102 identifies the time delay between the target device and the auxiliary device in receiving the user's voice command, and calculates the microphone array structure of the target device based on the time delay and the microphone array structure of the auxiliary device. For details, please refer to the relevant description of step S102 in the above method embodiment, which will not be repeated here.

[0099] The signal noise reduction module 103 performs noise reduction on the user's voice command based on the microphone array structure of the target device. For details, please refer to the relevant description of step S103 in the above method embodiment, which will not be repeated here.

[0100] A distributed microphone array noise reduction device provided in an embodiment of the present invention is used to execute a distributed microphone array noise reduction method provided in the above embodiment. Its implementation method and principle are the same. For details, please refer to the relevant description of the above method embodiment and will not be repeated here.

[0101] Through the collaborative cooperation of the above-mentioned components, the technical solution provided by the present application sets a control device and an auxiliary device of a known microphone array structure in a target area containing multiple smart devices. When the user issues a voice command, the signal strength information is calculated according to the voice command received by each smart device, and then each smart device sends the signal strength information to the control device. The control information finds the target device corresponding to the voice command with the strongest signal strength from the smart device according to the signal strength information, that is, accurately identifies the target device closest to the user. Then identify the time delay between the target device and the auxiliary device for receiving the user's voice command. According to the time delay and the known microphone array structure of the auxiliary device, a coordinate equation can be established to solve the spatial coordinates of each microphone device in the microphone array of the target device, and obtain the microphone array structure of the target device. Through the above steps, no matter how the position of the smart device changes, the microphone array structure of the smart device can be accurately obtained, and then the user's voice command can be denoised based on the microphone array structure of the target device combined with microphone noise reduction algorithms such as beamforming. Thereby solving the problem that the voice noise reduction algorithm is difficult to apply in the distributed microphone array scenario.

[0102] In addition, when the embodiment of the present invention performs noise reduction on the voice signal, the specific coordinate position of the sound source in space is first determined based on the microphone array structure of the target device, and then the user voice command signals with different time delays received by multiple microphones in the microphone array can be time-aligned according to the position of the sound source, so as to weight the sum of the user command voice signals received by each microphone. In the weighted summation process, based on the random characteristics of white noise, the noise will not be amplified, so that the noise ratio in the enhanced user voice command signal is greatly reduced, achieving the purpose of fast and simple noise reduction. In addition, based on the blocking matrix, multiple noise signals are extracted, and the enhanced user voice command is denoised using a filter, further improving the signal quality of the user voice command.

[0103] Figure 6 A control device according to an embodiment of the present invention is shown, the device includes a processor 901 and a memory 902, which can be connected via a bus or other means. Figure 6 The example of connecting through bus is taken in the following.

[0104] The processor 901 may be a central processing unit (CPU). The processor 901 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.

[0105] The memory 902 is a non-transitory computer-readable storage medium that can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as program instructions / modules corresponding to the methods in the above method embodiments. The processor 901 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory 902, that is, implementing the methods in the above method embodiments.

[0106] The memory 902 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required by at least one function; the data storage area may store data created by the processor 901, etc. In addition, the memory 902 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 902 may optionally include a memory remotely arranged relative to the processor 901, and these remote memories may be connected to the processor 901 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0107] One or more modules are stored in the memory 902 , and when executed by the processor 901 , the method in the above method embodiment is executed.

[0108] The specific details of the above-mentioned control device can be understood by referring to the corresponding related descriptions and effects in the above-mentioned method embodiment, and will not be repeated here.

[0109] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the implemented program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memory.

[0110] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A distributed microphone array noise reduction method, characterized in that: A control device applied to a target area, wherein the target area also includes a plurality of smart devices with microphone arrays, wherein the smart devices include at least one auxiliary device with a known microphone array structure, and the method includes: receiving signal strength information reported by each of the smart devices in response to a user voice command, and determining a target device to be interacted with by the user from each of the smart devices based on the signal strength information; Identifying a time delay between the target device and the auxiliary device in receiving the user voice command, and calculating a microphone array structure of the target device based on the time delay and the microphone array structure of the auxiliary device; A noise reduction operation is performed on the user voice command based on the microphone array structure of the target device.

2. The method according to claim 1, characterized in that: The step of determining a target device to be interacted with by the user from each of the smart devices based on each of the signal strength information includes: Compare the signal strength values ​​represented by each signal strength information, and obtain the target device information corresponding to the maximum signal strength value; A wake-up instruction is issued to the corresponding target device based on the target device information, so that the target device can interact with the user.

3. The method according to claim 1, characterized in that The performing noise reduction operation on the user voice command based on the microphone array structure of the target device includes: Calculating the sound source position of the user voice command based on the microphone array structure of the target device; Based on the microphone array structure of the target device and the sound source position, time-aligning the user voice command received by each microphone in the microphone array of the target device; The aligned user voice commands are weightedly summed to obtain the noise-reduced target user voice command.

4. The method according to claim 3, characterized in that: The calculating the sound source position of the user voice command based on the microphone array structure of the target device includes: Pairing microphones in the microphone array of the target device in pairs to obtain a plurality of microphone pairing combinations; The azimuth angle range and elevation angle range to be observed are divided into grids with preset intervals, and multiple angle observation grids are obtained by combining them in pairs; Based on the user voice command, traverse and calculate the angle spectrum function corresponding to each microphone pairing combination; The angle observation grid corresponding to the target angle spectrum function value in space is obtained to obtain the sound source position of the user voice command, wherein the target angle spectrum function value is the maximum function value in the angle spectrum function corresponding to each microphone pairing combination.

5. The method according to claim 4, characterized in that The calculating the sound source position of the user voice command based on the microphone array structure of the target device also includes: Dividing the user voice command into a plurality of equally spaced signal blocks, and dividing each signal block into a plurality of equally spaced signal frames; By traversing each signal block and generating a plurality of candidate angle observation grids, from the step of traversing and calculating the angle spectrum function corresponding to each microphone pairing combination based on the user voice command to the step of obtaining the angle observation grid corresponding to the target angle spectrum function value in space, a plurality of candidate angle observation grids are generated; Determining a second angle observation grid from the plurality of candidate angle observation grids based on data smoothing processing to obtain a sound source position of the user voice command; The target angle spectrum function value of each signal block is generated based on the maximum value of the angle spectrum function of the signal frame in the signal block.

6. The method according to claim 3, characterized in that The performing noise reduction operation on the user voice command based on the microphone array structure of the target device also includes: Inputting each user voice command after time alignment processing into a blocking matrix, and outputting a noise signal corresponding to each user voice command; The target user voice command and the multiple noise signals are simultaneously input into a preset adaptive filter, so as to use the multiple noise signals to perform a noise reduction operation on the user voice command after the weighted summation process, and obtain a second target user voice command.

7. The method according to claim 1, characterized in that The method further comprises: The control device and each smart device perform time synchronization.

8. A distributed microphone array noise reduction device, characterized in that: A control device applied to a target area, wherein the target area also includes a plurality of smart devices with microphone arrays, wherein the smart devices include at least one auxiliary device with a known microphone array structure, and the device includes: An interactive device determination module receives signal strength information reported by each of the smart devices in response to a user voice command, and determines a target device to be interacted with by the user from each of the smart devices based on the signal strength information; a microphone array estimation module, which identifies a time delay between the target device and the auxiliary device in receiving the user voice command, and calculates a microphone array structure of the target device based on the time delay and the microphone array structure of the auxiliary device; A signal noise reduction module performs noise reduction operations on the user voice command based on the microphone array structure of the target device.

9. A control device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Microphone-array speech-beam forming method as well as speech-signal processing device and system

    CN102324237A

  • Ultrasonic-assisted microphone array speech enhancement device

    CN102800325A