Microphone array voice noise reduction method and related equipment

By employing deep learning technology and a microphone spatial location-driven audio noise differential suppression mechanism, the problem of inconsistent noise reduction effects among multiple microphones in a microphone array is solved, achieving more efficient and consistent voice noise reduction and improving the noise reduction performance of voice-controlled robots.

CN121811907APending Publication Date: 2026-04-07UBTECH ROBOTICS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing speech noise reduction technologies cannot effectively solve the problem of consistent noise reduction effects across multiple microphones in microphone arrays. In particular, when faced with uneven audio energy distribution and different noise interference intensities at different microphones, significant differences in speech signals occur, affecting the noise reduction effect.

Method used

By employing deep learning technology combined with a microphone spatial location-driven audio noise differential suppression mechanism, the system performs signal framing processing, short-time Fourier transform, and amplitude spectrum adaptive correction on multiple speech signals. It utilizes a pre-trained speech denoising model for initial noise reduction and then performs adaptive correction based on the spatial noise suppression matrix of the microphone array to improve the noise reduction consistency of multiple speech signals.

Benefits of technology

It improves the consistency of speech noise reduction performance of microphone arrays in different environments, ensures the final noise-reduced speech quality, and solves the latency and resource consumption problems existing in traditional and multi-microphone deep learning solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811907A_ABST
    Figure CN121811907A_ABST
Patent Text Reader

Abstract

The invention provides a microphone array voice noise reduction method and related equipment, and relates to the technical field of voice noise reduction. According to the invention, after multiple paths of to-be-denoised voice signals collected by a target microphone array are converted into multiple paths of original frequency domain signals through signal framing processing and short-time Fourier transform processing, a voice denoising model is called to carry out preliminary denoising processing on the multiple paths of original frequency domain signals; according to a spatial noise suppression matrix of a target microphone array for multi-band sound, amplitude spectrum self-adaptive correction is carried out on multiple paths of preliminary noise reduction frequency domain signals obtained through preliminary noise reduction processing; and then, through short-time inverse Fourier transform processing and signal frame combination processing, multiple paths of target frequency domain signals obtained through amplitude spectrum correction processing are converted into multiple paths of target voice signals, so that voice noise reduction is realized through cooperation between a voice noise reduction mechanism based on a deep learning technology and an audio noise differential suppression mechanism based on microphone spatial position driving. And the noise reduction effect consistency of multiple paths of voice signals is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech noise reduction technology, and more specifically, to a microphone array speech noise reduction method and related equipment. Background Technology

[0002] With the continuous development of science and technology, voice control technology, as an important human-computer interaction mechanism, has been widely applied in the field of robotics, resulting in a variety of voice-controlled robots (such as home service robots and educational robots). Voice-controlled robots are often equipped with multiple individual microphones forming a complete microphone array to avoid missed voice signal acquisition through the coordinated operation of these microphones.

[0003] It is worth noting that because the actual working environment of voice-controlled robots is unpredictable, during the process of acquiring voice signals using microphone arrays, they will inevitably be affected by various low-frequency mechanical noises (with an audio frequency range concentrated in the 200-500Hz range), various high-frequency electronic noises (with an audio frequency range concentrated in the 10-16KHz range), and reverberation from multiple people speaking (with an audio frequency range concentrated in the 2-6KHz range). At the same time, due to the size limitations of the robot body and the need for a compact layout of the microphone array, the audio energy distribution of the same frequency band at different individual microphones is often different (i.e., the same frequency band causes different levels of interference to the voice acquisition at different individual microphones). This results in certain signal differences between the voice signals acquired simultaneously by the corresponding microphone array through multiple individual microphones, making it difficult for existing voice noise reduction technologies to meet the requirement of consistent noise reduction effects from multiple microphones. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a microphone array speech denoising method, a voice-controlled robot, and a readable storage medium, which can, on the basis of achieving the initial denoising effect of multiple speech signals using deep learning technology, introduce an audio noise differential suppression mechanism driven by microphone spatial position for audio adaptive correction, so as to effectively improve the consistency of the denoising effect of multiple speech signals and ensure the final denoised speech quality.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, this application provides a microphone array speech noise reduction method, the method comprising: The multiple audio signals to be denoised, collected by the target microphone array under the current working environment, are processed by signal frame segmentation and short-time Fourier transform to obtain multiple original frequency domain signals. The pre-trained speech denoising model is called to perform preliminary denoising on multiple original frequency domain signals, resulting in multiple preliminary denoised frequency domain signals. Based on the spatial noise suppression matrix of the target microphone array for multi-band sound, the amplitude spectrum adaptive correction is performed on the multiple preliminary noise reduction frequency domain signals to obtain multiple target frequency domain signals. The multi-target frequency domain signals are subjected to short-time inverse Fourier transform and signal framing processing respectively to obtain multi-target speech signals.

[0006] In an optional implementation, the step of calling a pre-trained speech denoising model to perform preliminary denoising processing on multiple original frequency domain signals to obtain multiple preliminary denoised frequency domain signals includes: A pre-trained speech denoising model is invoked to perform amplitude spectrum denoising on the first original frequency domain signal to obtain a first preliminary denoised frequency domain signal; wherein, the first original frequency domain signal is any one of the multiple original frequency domain signals; Based on the amplitude spectrum change data between the first preliminary denoising frequency domain signal and the first original frequency domain signal, amplitude spectrum synchronous denoising processing is performed on all second original frequency domain signals to obtain the corresponding second preliminary denoising frequency domain signal; wherein, each second original frequency domain signal is any one of the multiple original frequency domain signals other than the first original frequency domain signal.

[0007] In an optional implementation, the first original frequency domain signal includes multiple frames of first original spectrum data, and the first preliminary denoising frequency domain signal includes multiple frames of first denoising spectrum data. The step of calling a pre-trained speech denoising model to perform amplitude spectrum denoising on the first original frequency domain signal to obtain the first preliminary denoising frequency domain signal includes: For each frame of the first original spectrum data, sub-bands are divided according to the audio sampling rate value of the speech denoising model to obtain at least one first sub-band spectrum data; The speech denoising model is invoked to perform amplitude spectrum denoising on the at least one first sub-band spectrum data to obtain at least one second sub-band spectrum data. Subband merging is performed on the at least one second subband spectrum data to obtain a frame of first noise-reduced spectrum data.

[0008] In an optional implementation, the amplitude spectrum change data includes the amplitude spectrum ratio distribution data between the amplitude spectrum ratios of all first sub-band spectrum data and the corresponding second sub-band spectrum data involved in each of the multiple frames of first original spectrum data. Then, for each second original frequency domain signal, the step of performing amplitude spectrum synchronous noise reduction processing on the second original frequency domain signal to obtain the corresponding second preliminary noise-reduced frequency domain signal includes: For each frame of second original spectrum data included in the second original frequency domain signal, sub-bands are divided according to the audio sampling rate value of the speech denoising model to obtain at least one third sub-band spectrum data. For each third sub-band spectrum data, the amplitude spectrum is reduced according to the amplitude spectrum ratio distribution data related to the target first sub-band spectrum data to obtain the corresponding fourth sub-band spectrum data; wherein, the target first sub-band spectrum data and the third sub-band spectrum data are frequency domain aligned, and the first original spectrum data in which the target first sub-band spectrum data is located maintains the same frame order as the second original frequency domain signal; Subband merging is performed on the fourth subband spectrum data corresponding to each of the at least one second subband spectrum data to obtain a frame of second noise-reduced spectrum data including the second preliminary noise-reduced frequency domain signal.

[0009] In an optional implementation, the amplitude spectrum ratio distribution data related to the target first sub-band spectrum data includes the actual amplitude spectrum ratios before and after noise reduction for each frequency point within the corresponding sub-band frequency range. Therefore, the step of performing amplitude spectrum reduction processing on the third sub-band spectrum data according to the amplitude spectrum ratio distribution data related to the target first sub-band spectrum data to obtain the corresponding fourth sub-band spectrum data includes: For each frequency point involved in the third sub-band spectrum data, the theoretical amplitude spectrum value after noise reduction is calculated based on the actual amplitude spectrum value of the frequency point in the third sub-band spectrum data, according to the actual amplitude spectrum ratio before and after noise reduction corresponding to the frequency point. The theoretical amplitude spectrum value after noise reduction at this frequency point is taken as the target amplitude spectrum value at the corresponding fourth sub-band spectrum data of this frequency point.

[0010] In an optional implementation, the spatial noise suppression matrix includes noise suppression sub-matrices for each individual microphone in the target microphone array when facing multi-frequency sound. The step of performing amplitude spectrum adaptive correction on the multiple preliminary noise-reduced frequency domain signals according to the spatial noise suppression matrix of the target microphone array for multi-frequency sound to obtain multiple target frequency domain signals includes: For each preliminary noise reduction frequency domain signal, a target noise suppression sub-matrix matching the target single microphone to which the preliminary noise reduction frequency domain signal belongs is found in the spatial noise suppression matrix; The amplitude spectrum of each frame of denoised spectrum data involved in the initial denoised frequency domain signal is corrected according to the target noise suppression sub-matrix to obtain a frame of target spectrum data included in the corresponding target frequency domain signal.

[0011] In an optional implementation, the target noise suppression sub-matrix records the noise suppression gain coefficients of the corresponding target microphone unit when facing different audio frequency bands. The step of performing amplitude spectrum correction on each frame of the denoised frequency domain signal involved in the initial denoised frequency domain signal according to the target noise suppression sub-matrix to obtain a frame of target frequency domain signal includes: For each frame of noise-reduced spectral data, determine the target frequency band to which each frequency point belongs; For each target frequency band, according to the noise suppression gain coefficient corresponding to the target frequency band in the target noise suppression sub-matrix, the target amplitude spectrum value of each frequency point belonging to the target frequency band at the frame of noise reduction spectrum data is subjected to amplitude spectrum enhancement processing. The frame of denoised spectrum data after the amplitude spectrum enhancement operation for all the frequency points is completed is used as a frame of target spectrum data.

[0012] In an optional implementation, the method further includes: For each individual microphone in the target microphone array, the proportion of audio energy distribution of each audio frequency band in the current working environment at that individual microphone is statistically analyzed. Based on the proportion of audio energy distribution of each audio segment at the individual microphone, the noise suppression gain coefficient of each audio segment at the individual microphone is adaptively configured. The spatial noise suppression matrix is ​​obtained by constructing a matrix for the noise suppression gain coefficients of each individual microphone in the target microphone array and each of the audio frequency bands.

[0013] In an optional implementation, the speech noise reduction model is structured as a grouped temporal convolutional recurrent network; the noise suppression gain coefficient of the human voice frequency band at any single microphone is inversely correlated with its audio energy distribution ratio, while the noise suppression gain coefficient of the noise frequency band at any single microphone is positively correlated with its audio energy distribution ratio.

[0014] Secondly, this application provides a voice-controlled robot, including a processor, a memory, and a target microphone array composed of multiple individual microphones, wherein the target microphone array is used to collect external voice signals; The memory stores a computer program that can be executed by the processor, which can execute the computer program to implement the microphone array speech noise reduction method described in any of the foregoing embodiments for the target microphone array.

[0015] Thirdly, this application provides a readable storage medium storing a computer program thereon, which, when executed by a voice-controlled robot including a target microphone array, implements the microphone array speech noise reduction method described in any of the foregoing embodiments.

[0016] In this case, the beneficial effects of the embodiments of this application may include the following: This application performs signal framing and short-time Fourier transform processing on multiple audio signals to be denoised, collected by the target microphone array under the current working environment, to obtain multiple original frequency domain signals. Based on this, a pre-trained speech denoising model is called to perform preliminary denoising processing on the multiple original frequency domain signals, resulting in multiple preliminary denoised frequency domain signals. This utilizes deep learning technology to achieve preliminary denoising effects on multiple audio signals. Then, based on the spatial noise suppression matrix of the target microphone array for multi-frequency sound (which records the noise suppression gain coefficients of each individual microphone based on its spatial position when facing different sound bands), the amplitude spectrum adaptive correction is performed on the multiple preliminary denoised frequency domain signals to obtain multiple target frequency domain signals. Subsequently, the multiple target frequency domain signals are subjected to inverse short-time Fourier transform processing and signal framing processing to obtain multiple target speech signals. By introducing an audio noise differentiation suppression mechanism driven by microphone spatial position (which corresponds to the aforementioned spatial noise suppression matrix) for audio adaptive correction, the consistency of the denoising effect of multiple speech signals is improved, and the final denoised speech quality is ensured.

[0017] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A schematic diagram illustrating the composition of a voice-controlled robot provided in an embodiment of this application; Figure 2 This is one of the flowcharts illustrating the microphone array speech noise reduction method provided in the embodiments of this application; Figure 3 for Figure 2 A flowchart illustrating the execution process of step S220. Figure 4 for Figure 3A schematic diagram of the execution flow of the neutron step S221; Figure 5 for Figure 3 A schematic diagram of the execution flow of the neutron step S222; Figure 6 for Figure 2 A flowchart illustrating the execution process of step S230. Figure 7 This is the second schematic flowchart of the microphone array speech noise reduction method provided in the embodiments of this application.

[0020] Icons: 10-Voice-controlled robot; 11-Memory; 12-Processor; 13-Communication unit; 14-Target microphone array. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0022] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0023] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0024] In the description of this application, it should be understood that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are used only for the convenience of describing this application and simplifying the description, and are not intended to indicate or imply that the equipment or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0025] In the description of this application, it should also be noted that, unless otherwise expressly specified and limited, the terms "set up," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0026] Furthermore, it is understood in the description of this application that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Those skilled in the art will understand the specific meaning of the above terms in this application based on the specific circumstances.

[0027] Through diligent research, the applicant discovered that existing speech noise reduction technologies can be mainly categorized into the following three types: (1) Traditional single-microphone noise reduction methods (e.g., spectral subtraction, filtering, blind source separation, etc.) focus on achieving independent noise reduction effect of a single microphone. In essence, they do not consider the problem that different microphones will have different speech signals due to array layout, and cannot solve the problem of consistent noise reduction effect of multiple microphones.

[0028] (2) The multi-microphone full deep learning scheme (i.e., deep learning technology is used to process speech signals for each microphone separately) is essentially similar to the traditional single-microphone noise reduction method. It does not take into account the differences in speech signals between different microphones due to array layout, and cannot solve the problem of consistent noise reduction effect of multi-microphone. At the same time, because multiple speech signals call their respective speech noise reduction models in parallel for real-time inference, the noise reduction delay is basically over 500ms, which far exceeds the latency threshold requirement of real-time voice interaction (e.g., 150ms). In addition, this scheme requires a large amount of model storage resources and model inference resources, which often exceeds the hardware capacity of voice-controlled robots in actual applications.

[0029] (3) Multi-microphone local deep learning global multiplexing scheme (which selects a certain speech signal, uses deep learning technology to denoise the speech signal, and then uses the amplitude spectrum difference of the speech signal before and after denoising to the remaining speech signals in equal proportion or equal amplitude for denoising processing). Although it can effectively reduce the amount of model storage resources and model inference resources and ensure the real-time operation of multi-microphone speech denoising, it still does not consider the problem that different microphones will have different speech signals due to array layout. This often leads to over-denoising of the speech signals of the multiplexed deep learning results (for example, the microphones near the robot are mainly affected by high-frequency electronic noise, and their corresponding speech signals will lose high-frequency speech details after multiplexing deep learning results) or insufficient noise suppression (for example, the microphones near the user side are mainly affected by reverberation of multiple people talking and low-frequency mechanical noise, and their corresponding speech signals will have residual mechanical noise interference after multiplexing deep learning results). In essence, it cannot solve the problem of consistency of multi-microphone denoising effect.

[0030] To address this, the applicant developed a microphone array speech denoising method, a voice-controlled robot, and a readable storage medium. This method, based on the initial denoising effect achieved using deep learning technology for multiple speech signals, introduces an audio noise differential suppression mechanism driven by microphone spatial location for adaptive audio correction. This organic combination of the deep learning-based speech denoising mechanism and the microphone spatial location-driven audio noise differential suppression mechanism improves the consistency of denoising effects across multiple microphones (or multiple speech signals) and ensures the final denoised speech quality, effectively solving the technical problems of the three existing speech denoising technologies mentioned above.

[0031] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0032] Please refer to Figure 1 , Figure 1This is a schematic diagram of the composition of the voice-controlled robot 10 provided in this application embodiment. In this application embodiment, the voice-controlled robot 10, based on the synchronous acquisition of external voice signals using multiple individual microphones, considers the differences in deployment positions between the multiple individual microphones and the audio energy distribution perceived by each of the multiple individual microphones for different frequency bands in the same working environment. During the voice noise reduction process of multiple voice signals, it differentiates the microphones to perform noise differential suppression processing for different frequency bands, thereby effectively improving the consistency of the noise reduction effect of multiple voice signals and ensuring the final noise-reduced voice quality. The voice-controlled robot 10 can be, but is not limited to, educational robots, guidance robots, sweeping robots, etc., with voice interaction functions; each individual microphone of the voice-controlled robot 10 is responsible for the acquisition of one voice signal independently.

[0033] In this embodiment, the voice-controlled robot 10 may include a memory 11, a processor 12, a communication unit 13, and a target microphone array 14. The memory 11, processor 12, communication unit 13, and target microphone array 14 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components may be electrically connected via one or more communication buses or signal lines.

[0034] In this embodiment, the memory 11 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. The memory 11 is used to store computer programs, and the processor 12 can execute the computer programs accordingly after receiving execution instructions.

[0035] In this embodiment, the processor 12 can be an integrated circuit chip with signal processing capabilities. The processor 12 can be a general-purpose processor, including at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Network Processor (NP), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in this embodiment.

[0036] In this embodiment, the communication unit 13 is used to establish a communication connection between the voice-controlled robot 10 and other electronic devices via a network, and to send and receive data via the network, wherein the network includes wired communication networks and wireless communication networks. For example, the voice-controlled robot 10 can communicate with a user terminal through the communication unit 13 to obtain and store the interactive response actions configured by the user terminal for different voice commands. This allows the voice-controlled robot 10 to quickly execute an interactive response action (which may be, but is not limited to, robot wake-up actions, robot positioning actions, voice interaction actions, limb control actions, etc.) when it recognizes any voice command to be executed through its voice recognition function.

[0037] In this embodiment, the target microphone array 14 consists of multiple individual microphones used to collect external voice signals from the working environment of the voice-controlled robot 10. The individual microphones in the target microphone array 14 are positioned differently, and their perceived audio energy distribution for different frequency bands varies in the same working environment. Consequently, the interference intensity caused by the same frequency band sound to different individual microphones also varies.

[0038] Taking a four-microphone array educational robot as an example, due to the compact design of educational robots (usually with a diameter ≤30cm), four-microphone arrays generally adopt a four-corner equidistant layout (5-8cm spacing). This microphone array layout causes the two rear microphones (i.e., the rear left microphone and the rear right microphone) closer to the inside of the body to be mainly affected by high-frequency electronic noise, while the two front microphones (i.e., the front left microphone and the front right microphone) closer to the user are mainly affected by the reverberation of multiple people talking on the user side and low-frequency mechanical noise, resulting in the speech signal collected by the four microphones simultaneously exhibiting significant spatial heterogeneity.

[0039] In this embodiment, the voice-controlled robot 10 may pre-store a specific computer program related to the microphone array voice noise reduction function in the memory 11. By driving the processor 12 to execute the specific computer program, based on the initial noise reduction effect of multiple voice signals achieved by deep learning technology, an audio noise differential suppression mechanism driven by microphone spatial position is introduced for audio adaptive correction. This utilizes the organic cooperation between the voice noise reduction mechanism based on deep learning technology and the audio noise differential suppression mechanism driven by microphone spatial position to improve the consistency of noise reduction effect of multiple microphones (or multiple voice signals) and ensure the final noise-reduced voice quality.

[0040] Understandable, Figure 1 The block diagram shown is only a schematic diagram of one composition of the voice-controlled robot 10. The voice-controlled robot 10 may also include components such as... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.

[0041] In this application, to ensure that the aforementioned voice-controlled robot 10 can achieve preliminary noise reduction of multiple voice signals using deep learning technology, an audio noise differential suppression mechanism driven by microphone spatial position is introduced for adaptive audio correction. This effectively improves the consistency of noise reduction effects of multiple voice signals and ensures the final noise-reduced voice quality. The embodiments of this application provide a microphone array voice noise reduction method to achieve the aforementioned objective. The microphone array voice noise reduction method provided in this application will be described in detail below.

[0042] Please refer to Figure 2 , Figure 2 This is one of the flowcharts illustrating the microphone array speech noise reduction method provided in this application embodiment. In this application embodiment, Figure 2 The microphone array speech noise reduction method shown may include steps S210 to S240.

[0043] Step S210: Perform signal frame processing and short-time Fourier transform processing on the multiple audio signals to be denoised collected by the target microphone array under the current working environment to obtain multiple original frequency domain signals.

[0044] In this embodiment, after the voice-controlled robot 10 synchronously acquires multiple audio signals to be denoised using the multiple individual microphones included in the target microphone array 14 (each audio signal to be denoised corresponds to a single individual microphone), it performs frame processing on each of the multiple audio signals to be denoised according to the same fixed frame length and fixed frame shift. Then, it performs short-time Fourier transform processing on each frame of audio signal separated from each of the multiple audio signals to be denoised to obtain a frame of original spectrum data in the corresponding original frequency domain signal. Each audio signal to be denoised corresponds to a single original frequency domain signal; the fixed frame shift is less than the fixed frame length to ensure continuity and smoothness between adjacent frames of audio signals using overlapping segmentation technology.

[0045] In one embodiment of this example, the fixed frame length can be set to 512, and the fixed frame shift can be set to 256, so that two adjacent frames of speech signals have a 50% overlap. The window function selected for the corresponding short-time Fourier transform operation can be a square root Hanning window with the same length as the fixed frame length, so that the number of frequency points involved in any frame of original spectrum data obtained after short-time Fourier transform processing is 257 (i.e., 512 × 50% + 1).

[0046] Step S220: Call the pre-trained speech denoising model to perform preliminary denoising processing on the multiple original frequency domain signals to obtain multiple preliminary denoised frequency domain signals.

[0047] In this embodiment, the speech denoising model is a neural network model trained using deep learning technology, suitable for speech denoising processing. The model architecture of the speech denoising model can be, but is not limited to, grouped temporal convolutional recurrent networks, deep neural networks, recurrent neural networks, long short-term memory networks, etc. Each preliminary denoised frequency domain signal is obtained by denoising one original frequency domain signal. Specifically, the aforementioned multi-microphone full deep learning scheme can be used to call the speech denoising model for separate preliminary denoising processing of each of the multiple original frequency domain signals. However, to improve the real-time performance of multi-channel speech signal denoising operations and avoid excessive model storage and calculation resource consumption, the aforementioned multi-microphone local deep learning global reuse scheme can be used. Only one original frequency domain signal is directly processed using the speech denoising model for preliminary denoising, and based on the difference in amplitude spectrum before and after denoising of that original frequency domain signal, it is proportionally or equally multiplexed onto the remaining original frequency domain signals for denoising processing.

[0048] Alternatively, please refer to Figure 3 , Figure 3 yes Figure 2 A schematic diagram of the execution flow of step S220. In this embodiment of the application, step S220 may include sub-steps S221 and S222 to effectively improve the real-time performance of multi-channel speech signal noise reduction operation by utilizing the aforementioned multi-microphone local deep learning global multiplexing scheme, and avoid excessive model storage resource consumption and model inference resource consumption.

[0049] Sub-step S221: Call the speech denoising model to perform amplitude spectrum denoising on the first original frequency domain signal to obtain the first preliminary denoised frequency domain signal.

[0050] In this embodiment, the first original frequency domain signal is any one of the multiple original frequency domain signals. The first original frequency domain signal includes multiple frames of first original spectrum data, and the first preliminary denoising frequency domain signal includes multiple frames of first denoising spectrum data, wherein each frame of first denoising spectrum data is obtained by denoising one frame of first original spectrum data through the speech denoising model.

[0051] Alternatively, please refer to Figure 4 , Figure 4 yes Figure 3 A schematic diagram of the execution flow of sub-step S221. In this embodiment, considering the possibility that the audio sampling rate of the speech denoising model (e.g., 12kHz) may be inconsistent with the bandwidth of any frame of the first original spectrum data, sub-step S221 may include sub-steps S221a to S221c, so as to utilize the cooperation between the sub-band division mechanism and the denoising model multiple call mechanism to enable any frame of the first original spectrum data to achieve a preliminary denoising effect through deep learning technology, thereby obtaining a corresponding frame of first denoised spectrum data.

[0052] Sub-step S221a: For each frame of the first original spectrum data, sub-bands are divided according to the audio sampling rate value of the speech denoising model to obtain at least one first sub-band spectrum data.

[0053] In this embodiment, when the bandwidth of a frame of first original spectrum data is consistent with the audio sampling rate of the speech denoising model, the frame of first original spectrum data can be directly used as a first sub-band spectrum data. However, when the bandwidth of a frame of first original spectrum data is much larger than the audio sampling rate of the speech denoising model, the frame of first original spectrum data can be divided into multiple first sub-band spectrum data according to the audio sampling rate using an orthogonal mirror filter. For example, when the bandwidth of a single frame of first original spectrum data is 0~36KHz, and the audio sampling rate of the speech denoising model is 12KHz, the frame of first original spectrum data can be divided into first sub-band spectrum data of 0~12KHz, 12~24KHz, and 24~36KHz.

[0054] Sub-step S221b involves calling a speech denoising model to perform amplitude spectrum denoising on at least one first sub-band spectrum data to obtain at least one second sub-band spectrum data.

[0055] Each second sub-band spectrum data is obtained by denoising the amplitude spectrum of a first sub-band spectrum data.

[0056] Sub-step S221c involves sub-band merging of at least one second sub-band spectrum data to obtain a frame of first denoised spectrum data.

[0057] Therefore, by executing the above sub-steps S221a to S221c, this application can utilize the cooperation between the sub-band division mechanism and the noise reduction model multiple call mechanism to enable any frame of the first original spectrum data to achieve a preliminary noise reduction effect through deep learning technology, thereby obtaining the corresponding frame of the first noise-reduced spectrum data.

[0058] Sub-step S222: Based on the amplitude spectrum change data between the first preliminary denoising frequency domain signal and the first original frequency domain signal, perform amplitude spectrum synchronous denoising processing on all second original frequency domain signals to obtain the corresponding second preliminary denoising frequency domain signal.

[0059] In this embodiment, each second original frequency domain signal is any one of the multiple original frequency domain signals other than the first original frequency domain signal. Each second preliminary denoising frequency domain signal is obtained by amplitude spectrum synchronous denoising processing of one second original frequency domain signal. The multiple frames of second denoising spectrum data included in each second preliminary denoising frequency domain signal are obtained by denoising one frame of second original spectrum data in the corresponding second original frequency domain signal. The amplitude spectrum change data between the first preliminary denoising frequency domain signal and the first original frequency domain signal is used to describe the amplitude spectrum difference of the first preliminary denoising frequency domain signal before and after processing by the speech denoising model. It can be characterized by the amplitude spectrum ratio distribution data between each first sub-band spectrum data involved in each frame of first original spectrum data and the corresponding second sub-band spectrum data.

[0060] Alternatively, please refer to Figure 5 , Figure 5 yes Figure 3 The flowchart of sub-step S222 is shown below. In this embodiment, considering that any frame of the first original spectrum data may be divided into multiple first sub-band spectrum data during the speech denoising model call, and in order to ensure that the amplitude spectrum denoising result of the corresponding first original spectrum data can be reused in any second original frequency domain signal with the same frame order, sub-step S222 can achieve the aforementioned effect based on the sub-band division mechanism for each second original frequency domain signal. At this time, sub-step S222 may include sub-steps S222a to S222c.

[0061] Sub-step S222a: For each frame of second original spectrum data included in the second original frequency domain signal, sub-bands are divided according to the audio sampling rate value of the speech denoising model to obtain at least one third sub-band spectrum data.

[0062] Sub-step S222b: For each third sub-band spectrum data, according to the amplitude spectrum ratio distribution data related to the target first sub-band spectrum data, the amplitude spectrum of the third sub-band spectrum data is reduced to obtain the corresponding fourth sub-band spectrum data.

[0063] In this embodiment, for any third sub-band spectrum data, its corresponding target first sub-band spectrum data needs to be aligned with the frequency domain of the third sub-band spectrum data, and the first original spectrum data containing the target first sub-band spectrum data needs to maintain the same frame order as the second original frequency domain signal containing the third sub-band spectrum data. Wherein, the amplitude spectrum ratio distribution data related to the target first sub-band spectrum data includes the actual amplitude spectrum ratio before and after noise reduction for each frequency point within the corresponding sub-band frequency range. Therefore, sub-step S222b may include: For each frequency point involved in the third sub-band spectrum data, the theoretical amplitude spectrum value after noise reduction is calculated based on the actual amplitude spectrum value of the frequency point in the third sub-band spectrum data, according to the actual amplitude spectrum ratio before and after noise reduction corresponding to the frequency point. The theoretical amplitude spectrum value after noise reduction at this frequency point is taken as the target amplitude spectrum value at the corresponding fourth sub-band spectrum data of this frequency point.

[0064] In one embodiment of this example, the actual ratio between the actual amplitude spectrum value and the theoretical amplitude spectrum value after noise reduction at the same frequency point is consistent with the actual amplitude spectrum ratio before and after noise reduction corresponding to that frequency point.

[0065] In another embodiment of this example, the actual amplitude spectrum difference between the actual amplitude spectrum value and the theoretical amplitude spectrum value after noise reduction at the same frequency point can be calculated using the formula "[A-(A÷B)]", where "A" represents the actual amplitude spectrum value at the corresponding target first sub-band spectrum data at that frequency point, and "B" represents the actual amplitude spectrum ratio before and after noise reduction at that frequency point.

[0066] Sub-step S222c involves sub-band merging of the fourth sub-band spectrum data corresponding to at least one second sub-band spectrum data to obtain a frame of second noise-reduced spectrum data comprising the second preliminary noise-reduced frequency domain signal.

[0067] Therefore, by executing the above sub-steps S221 to S222, this application can effectively improve the real-time performance of multi-channel speech signal noise reduction operations by utilizing the aforementioned multi-microphone local deep learning global multiplexing scheme, and avoid excessive model storage resource consumption and model inference resource consumption.

[0068] Step S230: Based on the spatial noise suppression matrix of the target microphone array for multi-band sound, the amplitude spectrum adaptive correction is performed on the multiple preliminary noise reduction frequency domain signals to obtain multiple target frequency domain signals.

[0069] In this embodiment, the spatial noise suppression matrix describes the differentiated noise suppression intensity that each individual microphone needs to maintain when facing different frequency bands of sound. Its specific matrix element values ​​are strongly correlated with the audio energy distribution perceived by each individual microphone under the same working environment for different frequency bands of sound. The spatial noise suppression matrix may include noise suppression sub-matrices for each individual microphone in the target microphone array 14 when facing multiple frequency bands of sound. Each noise suppression sub-matrice records the noise suppression gain coefficient of the corresponding individual microphone when facing different audio frequencies (which positively characterizes the noise suppression intensity of the corresponding individual microphone for a specific audio frequency band). Based on this, the voice-controlled robot 10 can perform amplitude spectrum adaptive correction on multiple preliminary noise-reduced frequency domain signals based on the spatial noise suppression matrix to achieve adaptive audio correction for multiple microphones. This facilitates ensuring that multiple microphones can have a consistent and accurate noise reduction effect through the obtained multiple target frequency domain signals. At this time, any frame of target spectrum data included in each target frequency domain signal is obtained by processing the amplitude spectrum correction of a frame of noise-reduced spectrum data of the corresponding preliminary noise-reduced frequency domain signal.

[0070] Alternatively, please refer to Figure 6 , Figure 6 yes Figure 2 A schematic diagram of the execution flow of step S230. In this embodiment of the application, step S230 may include sub-steps S231 and S232 to introduce an audio noise differential suppression mechanism driven by microphone spatial position for adaptive audio correction, thereby effectively improving the consistency of noise reduction effect of multiple voice signals.

[0071] Sub-step S231: For each preliminary noise reduction frequency domain signal, find the target noise suppression sub-matrix in the spatial noise suppression matrix that matches the target single microphone to which the preliminary noise reduction frequency domain signal belongs.

[0072] Sub-step S232: According to the target noise suppression sub-matrix, the amplitude spectrum of each frame of noise reduction spectrum data involved in the initial noise reduction frequency domain signal is corrected to obtain a frame of target spectrum data included in the corresponding target frequency domain signal.

[0073] In this embodiment, the target noise suppression sub-matrix records the noise suppression gain coefficients of the corresponding target microphone when facing different audio frequency bands. Based on this, sub-step S232 may include: determining the target frequency band to which each frequency point belongs for all frequency points involved in each frame of noise-reduced spectrum data; for each target frequency band, according to the noise suppression gain coefficient corresponding to the target frequency band in the target noise suppression sub-matrix, performing amplitude spectrum enhancement processing on the target amplitude spectrum value of each frequency point belonging to the target frequency band in the frame of noise-reduced spectrum data (which is achieved by multiplying the noise suppression gain coefficient of the target frequency band with the target amplitude spectrum value of a single frequency point within the target frequency band); and taking the frame of noise-reduced spectrum data after completing the amplitude spectrum enhancement operation for all frequency points as a frame of target spectrum data.

[0074] Therefore, by executing the above sub-steps S231 and S232, this application can differentiate the microphones and perform noise suppression processing for different frequency bands during the speech denoising process of multiple speech signals, so as to effectively improve the consistency of the denoising effect of multiple speech signals and ensure the final denoised speech quality.

[0075] Step S240: Perform short-time inverse Fourier transform and signal framing processing on the multiple target frequency domain signals respectively to obtain multiple target speech signals.

[0076] In this embodiment, the short-time Fourier transform operation involved in step S240 and the short-time Fourier transform operation involved in step S210 use the same window function. The signal framing operation involved in step S240 and the signal framing operation involved in step S210 both use the same fixed frame length and fixed frame shift.

[0077] Therefore, by performing the above steps S210 to S240, this application can, on the basis of achieving the initial noise reduction effect of multiple speech signals using deep learning technology, introduce an audio noise differential suppression mechanism driven by microphone spatial position to perform audio adaptive correction, so as to effectively improve the consistency of noise reduction effect of multiple speech signals and ensure the final noise-reduced speech quality.

[0078] Alternatively, please refer to Figure 7 , Figure 7 This is the second schematic flowchart of the microphone array speech noise reduction method provided in this application embodiment. In this application embodiment, with Figure 2 Compared to the microphone array speech noise reduction method shown, Figure 7The microphone array speech noise reduction method shown may further include steps S250 to S270, in order to comprehensively consider the audio energy distribution of each individual microphone in the target microphone array 14 under the same working environment for different frequency bands of sound, and from the perspective of the difference in the spatial layout of the microphones, ensure that the final configured spatial noise suppression matrix can distinguish the microphones to achieve differentiated noise suppression effects for different frequency bands of sound under the same working environment, so as to ensure the consistency of the noise reduction effect of multiple speech signals and the quality of the noise-reduced speech.

[0079] Step S250: For each individual microphone in the target microphone array, calculate the proportion of audio energy distribution of each audio frequency band in the current working environment at that individual microphone.

[0080] In this embodiment, the proportion of audio energy distribution of different audio bands in the same single microphone can positively reflect the overall influence of the corresponding audio band in the speech acquisition process of that single microphone. Taking the "educational robot with a four-microphone array" described above as an example, the audio energy distribution of "multi-person speaking reverberation" at the front left / right microphones accounts for approximately 65% ​​to 75%, "low-frequency mechanical noise" accounts for approximately 20% to 24%, and "high-frequency electronic noise" has an audio energy below -45dB at the front left / right microphones. In other words, the front left / right microphones mainly collect multi-person speaking reverberation and low-frequency mechanical noise from the user side during voice acquisition. Meanwhile, the audio energy distribution of "multi-person speaking reverberation" at the rear left / right microphones accounts for approximately 40% to 50%, and "high-frequency electronic noise" accounts for approximately 45% to 60%. Furthermore, the audio energy of "high-frequency electronic noise" at the rear left / right microphones is above -35dB. In other words, the rear left / right microphones mainly collect multi-person speaking reverberation and high-frequency electronic noise during voice acquisition.

[0081] Step S260: Based on the proportion of audio energy distribution of each audio segment at the individual microphone, adaptively configure the noise suppression gain coefficient of each audio segment at the individual microphone.

[0082] In this embodiment, each individual microphone in the target microphone array 14 can use the same gain coefficient range (e.g., 0.8~1.4) for noise suppression gain coefficient configuration when facing different audio frequency bands, or it can use different gain coefficient ranges (e.g., the gain coefficient range for the 2-6KHz human voice band is 0.85~1.1, the gain coefficient range for the 200-500Hz low-frequency mechanical noise is 1~1.3, and the gain coefficient range for the 200-500Hz low-frequency mechanical noise is 1~1.4) for noise suppression gain coefficient configuration.

[0083] Among them, the noise suppression gain coefficient of the human voice band at any single microphone is inversely correlated with its audio energy distribution ratio (i.e., the larger the audio energy distribution ratio of the human voice band, the smaller the actual noise suppression gain coefficient configured within the corresponding gain coefficient range; for example, the noise suppression gain coefficient of the human voice band at the front left / right microphone can be configured to 0.9 to protect the details of the human voice, and the noise suppression gain coefficient of the human voice band at the rear left / right microphone can be configured to 0.95 to moderately protect the human voice), while the noise suppression gain coefficient of the noise band (including the low-frequency mechanical noise band and the high-frequency electronic noise band) at any single microphone is positively correlated with its audio energy distribution ratio (i.e., the low-frequency mechanical noise band and the high-frequency electronic noise band...). The larger the proportion of audio energy distribution in each frequency band, the larger the actual noise suppression gain coefficient configured within the corresponding gain coefficient range. For example, the noise suppression gain coefficient of the low-frequency mechanical noise band at the front left / right microphone can be configured to 1.2 to enhance the mechanical noise suppression capability, the noise suppression gain coefficient of the high-frequency electronic noise band at the front left / right microphone can be configured to 1 (i.e., there is no need to over-suppress the high-frequency electronic noise), the noise suppression gain coefficient of the high-frequency electronic noise band at the rear left / right microphone can be configured to 1.3 to enhance the high-frequency electronic noise suppression capability, and the noise suppression gain coefficient of the low-frequency mechanical noise band at the rear left / right microphone can be configured to 1 (i.e., there is no need to over-suppress the low-frequency mechanical noise)).

[0084] Step S270: Construct a matrix for the noise suppression gain coefficients of each individual microphone in the target microphone array and each audio frequency band to obtain the spatial noise suppression matrix.

[0085] Therefore, by executing the above steps S250 to S270, this application comprehensively considers the audio energy distribution of each individual microphone in the target microphone array 14 under the same working environment for different frequency bands of sound. From the perspective of the differences in the spatial layout of the microphones, it ensures that the final configured spatial noise suppression matrix can distinguish the microphones and achieve differentiated noise suppression effects for different frequency bands of sound under the same working environment, thereby ensuring the consistency of noise reduction effect and the quality of noise-reduced speech signals for multiple voice signals.

[0086] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of the apparatus, methods, and computer program products according to embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0087] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part. If the function is implemented as a software functional module and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium and includes several instructions to cause a voice-controlled robot 10 equipped with a target microphone array 14 to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, and other media capable of storing program code.

[0088] The above descriptions are merely various embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A microphone array speech noise reduction method, characterized in that, The method includes: The multiple audio signals to be denoised, collected by the target microphone array under the current working environment, are processed by signal frame segmentation and short-time Fourier transform to obtain multiple original frequency domain signals. The pre-trained speech denoising model is called to perform preliminary denoising on multiple original frequency domain signals, resulting in multiple preliminary denoised frequency domain signals. Based on the spatial noise suppression matrix of the target microphone array for multi-band sound, the amplitude spectrum adaptive correction is performed on the multiple preliminary noise reduction frequency domain signals to obtain multiple target frequency domain signals. The multi-target frequency domain signals are subjected to short-time inverse Fourier transform and signal framing processing respectively to obtain multi-target speech signals.

2. The method according to claim 1, characterized in that, The step of calling a pre-trained speech denoising model to perform preliminary denoising processing on multiple original frequency domain signals to obtain multiple preliminary denoised frequency domain signals includes: The speech denoising model is invoked to perform amplitude spectrum denoising on the first original frequency domain signal to obtain a first preliminary denoised frequency domain signal; wherein, the first original frequency domain signal is any one of the multiple original frequency domain signals; Based on the amplitude spectrum change data between the first preliminary denoising frequency domain signal and the first original frequency domain signal, amplitude spectrum synchronous denoising processing is performed on all second original frequency domain signals to obtain the corresponding second preliminary denoising frequency domain signal; wherein, each second original frequency domain signal is any one of the multiple original frequency domain signals other than the first original frequency domain signal.

3. The method according to claim 2, characterized in that, The first original frequency domain signal includes multiple frames of first original spectrum data, and the first preliminary denoising frequency domain signal includes multiple frames of first denoising spectrum data. The step of calling a pre-trained speech denoising model to perform amplitude spectrum denoising on the first original frequency domain signal to obtain the first preliminary denoising frequency domain signal includes: For each frame of the first original spectrum data, sub-bands are divided according to the audio sampling rate value of the speech denoising model to obtain at least one first sub-band spectrum data; The speech denoising model is invoked to perform amplitude spectrum denoising on the at least one first sub-band spectrum data to obtain at least one second sub-band spectrum data. Subband merging is performed on the at least one second subband spectrum data to obtain a frame of first noise-reduced spectrum data.

4. The method according to claim 3, characterized in that, The amplitude spectrum change data includes the amplitude spectrum ratio distribution data between the first sub-band spectrum data and the corresponding second sub-band spectrum data involved in each of the multiple frames of first original spectrum data. The step of performing amplitude spectrum synchronous noise reduction processing on each second original frequency domain signal to obtain the corresponding second preliminary noise-reduced frequency domain signal includes: For each frame of second original spectrum data included in the second original frequency domain signal, sub-bands are divided according to the audio sampling rate value of the speech denoising model to obtain at least one third sub-band spectrum data. For each third sub-band spectrum data, the amplitude spectrum is reduced according to the amplitude spectrum ratio distribution data related to the target first sub-band spectrum data to obtain the corresponding fourth sub-band spectrum data; wherein, the target first sub-band spectrum data and the third sub-band spectrum data are frequency domain aligned, and the first original spectrum data in which the target first sub-band spectrum data is located maintains the same frame order as the second original frequency domain signal; Subband merging is performed on the fourth subband spectrum data corresponding to each of the at least one second subband spectrum data to obtain a frame of second noise-reduced spectrum data including the second preliminary noise-reduced frequency domain signal.

5. The method according to claim 4, characterized in that, The amplitude spectrum ratio distribution data related to the target first sub-band spectrum data includes the actual amplitude spectrum ratios before and after noise reduction for each frequency point within the corresponding sub-band frequency range. The step of performing amplitude spectrum reduction processing on the third sub-band spectrum data according to the amplitude spectrum ratio distribution data related to the target first sub-band spectrum data to obtain the corresponding fourth sub-band spectrum data includes: For each frequency point involved in the third sub-band spectrum data, the theoretical amplitude spectrum value after noise reduction is calculated based on the actual amplitude spectrum value of the frequency point in the third sub-band spectrum data, according to the actual amplitude spectrum ratio before and after noise reduction corresponding to the frequency point. The theoretical amplitude spectrum value after noise reduction at this frequency point is taken as the target amplitude spectrum value at the corresponding fourth sub-band spectrum data of this frequency point.

6. The method according to claim 1, characterized in that, The spatial noise suppression matrix includes noise suppression sub-matrices for each individual microphone in the target microphone array when facing multi-frequency sound. The step of performing amplitude spectrum adaptive correction on the multiple preliminary noise-reduced frequency domain signals according to the spatial noise suppression matrix of the target microphone array for multi-frequency sound to obtain multiple target frequency domain signals includes: For each preliminary noise reduction frequency domain signal, a target noise suppression sub-matrix matching the target single microphone to which the preliminary noise reduction frequency domain signal belongs is found in the spatial noise suppression matrix; The amplitude spectrum of each frame of denoised spectrum data involved in the initial denoised frequency domain signal is corrected according to the target noise suppression sub-matrix to obtain a frame of target spectrum data included in the corresponding target frequency domain signal.

7. The method according to claim 6, characterized in that, The target noise suppression sub-matrix records the noise suppression gain coefficients of the corresponding target microphone unit when facing different audio frequency bands. The step of performing amplitude spectrum correction on each frame of the denoised frequency domain signal involved in the initial denoised frequency domain signal according to the target noise suppression sub-matrix to obtain a frame of target frequency domain signal includes: For each frame of noise-reduced spectral data, determine the target frequency band to which each frequency point belongs; For each target frequency band, according to the noise suppression gain coefficient corresponding to the target frequency band in the target noise suppression sub-matrix, the target amplitude spectrum value of each frequency point belonging to the target frequency band at the frame of noise reduction spectrum data is subjected to amplitude spectrum enhancement processing. The frame of denoised spectrum data after the amplitude spectrum enhancement operation for all the frequency points is completed is used as a frame of target spectrum data.

8. The method according to any one of claims 1-7, characterized in that, The method further includes: For each individual microphone in the target microphone array, the proportion of audio energy distribution of each audio frequency band in the current working environment at that individual microphone is statistically analyzed. Based on the proportion of audio energy distribution of each audio segment at the individual microphone, the noise suppression gain coefficient of each audio segment at the individual microphone is adaptively configured. The spatial noise suppression matrix is ​​obtained by constructing a matrix for the noise suppression gain coefficients of each individual microphone in the target microphone array and each of the audio frequency bands.

9. A voice-controlled robot, characterized in that, It includes a processor, a memory, and a target microphone array consisting of multiple individual microphones, wherein the target microphone array is used to acquire external speech signals; The memory stores a computer program that can be executed by the processor to implement the microphone array speech noise reduction method according to any one of claims 1-8 for the target microphone array.

10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a voice-controlled robot including a target microphone array, it implements the microphone array speech noise reduction method according to any one of claims 1-8.