Voice input noise reduction processing method based on deep learning

By monitoring changes in microphone pointing angle and sensitivity in real time, stable microphones are selected and frequency domain convolution operators are used for frequency band division and noise reduction. This solves the noise recognition error problem caused by microphone offset and sensitivity changes in existing technologies, and improves speech clarity and recognition accuracy.

CN121884844APending Publication Date: 2026-04-17SHANGHAI MAIJUN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI MAIJUN TECHNOLOGY CO LTD
Filing Date
2026-01-23
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing voice input systems lack real-time detection and evaluation of microphone pointing angle deviation and dynamic changes in sensitivity, resulting in increased noise recognition errors, reduced accuracy of speech spectrum division, and decreased speech clarity and recognition accuracy.

Method used

By monitoring changes in microphone pointing angle and sensitivity in real time, stable microphones are selected. The frequency domain convolution operator is used to extract the spectral edge features of the speech signal, generate an adaptive noise reduction weight index, and dynamically mark and denoise the frequency band.

Benefits of technology

It improves the noise reduction effect of speech signal processing, optimizes speech clarity and recognition reliability, and enhances communication quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884844A_ABST
    Figure CN121884844A_ABST
Patent Text Reader

Abstract

The invention discloses a voice input noise reduction processing method based on deep learning, relates to the technical field of voice noise reduction, is used for solving the problem that noise recognition errors are increased, and aims to solve the problem that the noise recognition errors are increased by monitoring voice acquisition microphones, acquiring pointing angles and sensitivity variations of the microphones and calculating pickup angle offset amplitudes according to the pointing angles and the sensitivity variations. The method comprises the following steps: comprehensively evaluating a microphone pickup state, screening a stable microphone and calling a voice signal segment covered by the microphone, detecting spectrum edge information of the signal segment by using a time-frequency domain convolution operator, carrying out frequency band division based on the edge information, and obtaining a noise power ratio and voice signal energy intensity of each frequency band. Evaluating a frequency band noise change trend according to the noise power ratio, generating a noise reduction weight index in combination with the voice signal energy intensity, and marking the frequency band according to the index; by accurately screening the stable microphone and effectively evaluating the frequency band noise characteristics, the noise reduction effect of voice signal processing is improved, and the voice recognition and communication quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech noise reduction technology, and more specifically, to a speech input noise reduction processing method based on deep learning. Background Technology

[0002] With the rapid development of AI voice interaction, remote conferencing systems and voice recognition technology, voice signal acquisition and noise reduction processing have become important aspects affecting the accuracy of voice recognition and communication quality. Existing voice input systems typically use multi-microphone arrays to acquire external voice signals and suppress noise signals through beamforming, adaptive filtering or spectral subtraction.

[0003] The existing technology has the following shortcomings: Currently, existing technologies rely solely on the static calibration parameters of fixed microphone arrays during speech input noise reduction processing. They lack a real-time detection and evaluation mechanism for microphone pointing angle deviation and dynamic changes in sensitivity. This makes it impossible to accurately select microphones with stable pickup states and their effective speech signal segments, leading to increased noise recognition errors and reduced accuracy in speech spectrum division. Consequently, speech clarity and recognition accuracy decrease significantly. Therefore, a speech input noise reduction processing method based on deep learning is proposed.

[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a speech input noise reduction processing method based on deep learning. This method utilizes real-time monitoring of microphone pointing angle and sensitivity changes to select stable microphones, and combines time-frequency domain convolution operators to extract spectral edge features of the speech signal. Based on the noise power ratio and speech energy intensity, an adaptive noise reduction weight index is generated to dynamically label and denoise the frequency band, thereby solving the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a speech input noise reduction processing method based on deep learning, comprising the following steps: Step S1: Monitor the microphones used for voice acquisition, obtain the pointing angle and sensitivity change of each microphone, and calculate the pickup angle offset of each microphone based on the pointing angle. Step S2: Evaluate the pickup status of each microphone by combining the pickup angle offset amplitude and sensitivity change, filter the microphones according to the pickup status to obtain stable microphones, and retrieve the speech signal segment covered by the stable microphones. Step S3: Detect the spectral edge information of the speech signal segment using the time-frequency domain convolution operator, divide the speech signal segment into frequency bands based on the spectral edge information, and obtain the noise power ratio and speech signal energy intensity of each divided frequency band; Step S4: Evaluate the noise change trend of each divided frequency band based on the noise power ratio, generate the noise reduction weight index of each divided frequency band in combination with the speech signal energy intensity, and mark the divided frequency bands based on the noise reduction weight index.

[0007] In a preferred embodiment, in step S1, the pointing angle of each microphone is obtained by a rotary encoder mounted on the shaft of each microphone. The received sound wave signal is converted into an electrical signal by the built-in capacitive sensor of each microphone, which serves as the input signal for each microphone. The external sound wave signal is converted into an electrical signal by the dynamic sensor built into each microphone, which serves as the response signal for each microphone. The ratio of the response signal of each microphone to the input signal is used as the sensitivity of each microphone.

[0008] In a preferred embodiment, in step S1, the absolute value of the difference between the sensitivity of each microphone and the preset standard sensitivity is taken as the sensitivity change of each microphone. The absolute value of the difference between the pointing angle of each microphone and the preset reference angle is used as the pickup angle offset of each microphone.

[0009] In a preferred embodiment, in step S2, multiple sets of pickup angle offset amplitude and sensitivity change amount are obtained, and the pickup angle offset amplitude dataset and sensitivity change amount dataset are integrated. A multilayer perceptron model was constructed by combining the datasets of pickup angle offset amplitude and sensitivity change to analyze the pickup status of each microphone.

[0010] In a preferred embodiment, in step S2, a first stable state threshold is preset to be greater than a second stable state threshold. If the pickup state of each microphone is greater than the preset second stable state threshold and less than the preset first stable state threshold, then the microphone is determined to be a stable microphone. Otherwise, the microphone is determined to be an unstable microphone.

[0011] In a preferred embodiment, in step S3, the speech signal segment covered by the stable microphone is divided into multiple time windows, and a short-time Fourier transform is performed on each window to obtain a two-dimensional time-frequency diagram. The spectral edge information of speech signal segments is obtained by processing the two-dimensional time-frequency graph using a time-frequency domain convolution operator.

[0012] In a preferred embodiment, in step S3, if the spectral edge information of the speech signal segment is less than a preset frequency band range, the speech signal segment covered by the stable microphone is divided into a low frequency band. If the spectral edge information of a speech signal segment is within the preset frequency band, then the speech signal segment covered by the stable microphone will be divided into the mid-frequency band. If the spectral edge information of a speech signal segment is greater than the preset frequency band range, then the speech signal segment covered by the stable microphone will be divided into a high-frequency band.

[0013] In a preferred embodiment, in step S3, the noise-reduced power of each divided frequency band is obtained by spectral subtraction; The power spectrum of each frequency band is obtained by squaring each point in the two-dimensional time-frequency graph. The total power of each frequency band is obtained by summing the power of each time window in each frequency band. The noise power percentage of each frequency band is obtained by subtracting 1 from the ratio of the noise-reduced power of each frequency band to the total power of each frequency band and taking the absolute value. The energy intensity of the speech signal in each frequency band is obtained by summing the points of the power spectrum.

[0014] In a preferred embodiment, in step S4, the difference in the noise power ratio of adjacent time windows of each frequency band is taken as the noise change of each frequency band. The noise variation trend of each frequency band is obtained by averaging the noise variation of each frequency band. The noise variation trend and speech signal energy intensity of each frequency band are standardized to obtain the noise variation factor and speech intensity factor. The noise reduction weight index for each frequency band is calculated by combining the noise variation factor and the speech intensity factor.

[0015] In a preferred embodiment, in step S4, if the noise reduction weight index of the divided frequency band is greater than or equal to a preset noise reduction threshold, it is determined that the divided frequency band is marked; if the noise reduction weight index of the divided frequency band is less than the preset noise reduction threshold, it is determined that the divided frequency band is not marked.

[0016] The technical effects and advantages of this invention are as follows: This invention monitors the microphones used for voice acquisition, obtaining the changes in the pointing angle and sensitivity of each microphone. Based on the pointing angle, it calculates the pickup angle offset amplitude. By comprehensively considering the pickup angle offset amplitude and sensitivity changes, it evaluates the microphone's pickup status, selects stable microphones, and retrieves the speech signal segments they cover. Next, it uses a time-frequency domain convolution operator to detect the spectral edge information of the signal segments, divides frequency bands based on the edge information, and obtains the noise power ratio and speech signal energy intensity of each frequency band. Finally, it evaluates the frequency band noise change trend based on the noise power ratio, generates a noise reduction weight index based on the speech signal energy intensity, and marks the frequency bands according to this index. By accurately selecting stable microphones and effectively evaluating frequency band noise characteristics, it improves the noise reduction effect of speech signal processing, thereby optimizing speech clarity and reliability, and enhancing speech recognition and communication quality. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the implementation of the deep learning-based speech input noise reduction method of the present invention.

[0018] Figure 2 This is a schematic diagram illustrating the steps of the deep learning-based speech input noise reduction method of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] This invention monitors the microphones used for voice acquisition, obtaining the changes in the pointing angle and sensitivity of each microphone. Based on the pointing angle, it calculates the pickup angle offset amplitude. By comprehensively evaluating the pickup angle offset amplitude and sensitivity changes, it assesses the microphone's pickup status, selects stable microphones, and retrieves the speech signal segments they cover. Next, it uses a time-frequency domain convolution operator to detect the spectral edge information of the signal segments, divides frequency bands based on the edge information, and obtains the noise power ratio and speech signal energy intensity of each frequency band. Finally, it assesses the frequency band noise change trend based on the noise power ratio, generates a noise reduction weight index based on the speech signal energy intensity, and marks the frequency bands according to this index. By accurately selecting stable microphones and effectively assessing frequency band noise characteristics, it improves the noise reduction effect of speech signal processing.

[0021] Example 1 Deep learning-based speech input noise reduction methods, such as Figures 1 to 2 As shown, it includes the following steps: Step S1: Monitor the microphones used for voice acquisition, obtain the pointing angle and sensitivity change of each microphone, and calculate the pickup angle offset of each microphone based on the pointing angle. Step S2: Evaluate the pickup status of each microphone by combining the pickup angle offset amplitude and sensitivity change, filter the microphones according to the pickup status to obtain stable microphones, and retrieve the speech signal segment covered by the stable microphones. Step S3: Detect the spectral edge information of the speech signal segment using the time-frequency domain convolution operator, divide the speech signal segment into frequency bands based on the spectral edge information, and obtain the noise power ratio and speech signal energy intensity of each divided frequency band; Step S4: Evaluate the noise change trend of each divided frequency band based on the noise power ratio, generate the noise reduction weight index of each divided frequency band in combination with the speech signal energy intensity, and mark the divided frequency bands based on the noise reduction weight index.

[0022] The specific implementation is as follows: In step S1, the pointing angle of each microphone is obtained by a rotary encoder installed on the shaft of each microphone; The received sound wave signal is converted into an electrical signal by the built-in capacitive sensor of each microphone, which serves as the input signal for each microphone. The external sound wave signal is converted into an electrical signal by the dynamic sensor built into each microphone, which serves as the response signal for each microphone. The ratio of the response signal of each microphone to the input signal is used as the sensitivity of each microphone. The absolute value of the difference between the sensitivity of each microphone and the preset standard sensitivity is taken as the sensitivity change of each microphone. The absolute value of the difference between the pointing angle of each microphone and the preset reference angle is used as the pickup angle offset of each microphone.

[0023] It needs to be explained that a rotary encoder is a sensor that measures angle or rotational displacement. It converts rotational mechanical motion into digital or pulse signals, providing information such as angle, position, and speed, which is used to determine the pointing angle of each microphone. A capacitive sensor is a sensor that detects physical quantities based on the principle of capacitance change. It measures the capacitance change caused by sound waves and converts the sound wave signal into an electrical signal. A moving coil sensor is a sensor that converts sound wave signals into electrical signals. Its basic principle is to use the vibration caused by sound waves to generate an electrical signal. The input signal refers to the external sound wave signal received by the microphone. This is the physical form of the sound received by the microphone, and it is usually converted into an electrical signal by the microphone's capacitive or moving coil sensor. The input signal reflects the actual sound fluctuations in the environment. The response signal is the electrical signal output by the microphone in response to changes in the input signal. It represents the microphone's actual reaction to external sound waves and includes the microphone's inherent characteristics, such as frequency response, sensitivity, and noise interference. The preset standard sensitivity is a sensitivity value obtained through calibration testing of the microphone under specific standard test conditions. It can be obtained by statistically analyzing historical sensitivity data of microphones of the same model during factory inspection, taking the average or median sensitivity of that model as the preset standard sensitivity. The preset reference angle is a reference angle used to describe the ideal pickup direction of the microphone. A specific direction is usually chosen as a standard to indicate the direction the microphone should point when working normally; this is usually directly in front of the microphone, which is the microphone's strongest pickup direction.

[0024] In step S2, multiple sets of pickup angle offset amplitude and sensitivity change amount are obtained, and the pickup angle offset amplitude dataset and sensitivity change amount dataset are integrated. A multilayer perceptron model was constructed to analyze the pickup status of each microphone by combining the datasets of pickup angle offset amplitude and sensitivity change. The multilayer perceptron model consists of three layers: an input layer, a hidden layer, and an output layer. The input layer receives input features and converts them into easily manipulated data. The hidden layer performs calculations on the converted data, and the output layer outputs the calculation results. The specific steps are as follows: Input data: The pickup angle offset amplitude dataset and the sensitivity change dataset are used as input features and fed into the input layer. The input layer normalizes the data in the pickup angle offset amplitude dataset and the sensitivity change dataset and then feeds them into the hidden layer. Initialization parameters: In a multilayer perceptron model, the initial weights and bias parameters from the input layer to the hidden layer are set; Setting up a hidden layer processing algorithm: The pickup angle offset amplitude dataset and sensitivity change dataset are processed by setting a hidden layer activation function. Taking a relatively simple weighted function as an example, the activation function can be constructed as follows: ,in The output of the hidden layer, This represents the normalized value of the microphone corresponding to the pickup angle offset amplitude dataset. The normalized values ​​for the microphone corresponding to the sensitivity variation dataset are: and These are two initial weights set between the input layer and the hidden layer. These are the bias parameters set between the output layer and the hidden layer. Set the output layer processing algorithm: Set the activation function of the output layer: ,in The output layer outputs the results. This is the initial pickup state value; it can be set to the hidden layer output result calculated by averaging the pickup angle offset amplitude dataset and the sensitivity change dataset. Propagation: The multilayer perceptron passes the input features sequentially through the input layer, hidden layer, and output layer, and finally the output layer calculates the output result as the sound pickup state of each microphone; Calculate the pickup status value: Step 1: Calculate the loss function by comparing the output of the multilayer perceptron model with the preset actual labels. The mean squared error of the outputs of all output layers can be calculated using the loss function. ,in The mean squared error is denoted by n, where n is the number of data points in the dataset of pickup angle offset amplitude or sensitivity change. The first one randomly selected from the pickup angle offset amplitude dataset The output layer outputs the result obtained by passing the normalized result of the pickup angle offset amplitude and the normalized result of the sensitivity change at the corresponding sampling time in the dataset. For the preset first One actual label; Step 2: Calculate the gradients of the hidden layer and input layer's influence on the loss function. Use the partial derivatives of the hidden layer's weights and bias parameters with respect to the loss function as the gradients. The formulas for calculating the gradients of the weights and bias parameters are as follows: , ,in and These are the gradients of the weights and the gradients of the bias parameters in the activation function of the hidden layer, respectively. Step 3: Set the learning rate, which can be referenced from the values ​​used in similar language acquisition optimization scenarios for multilayer perceptrons; adjust the weights and bias parameters using the weight gradient, bias parameter gradient, and learning rate. For example, update the weights and bias parameters using the product of the learning rate and the corresponding weight gradient or bias parameter gradient: , ,in, For learning rate, For the adjusted weights or , To adjust the bias parameters; repeat the above steps until the model converges, take the final output of the hidden layer as the pickup state, and replace the initial pickup state value with the calculated pickup state value; Output results: Output results of the output layer after the model converges, and use the output results of the output layer as the sound pickup status of each microphone; It should be noted that the weights and bias parameters in the activation function of the hidden layer are randomly assigned and initialized. The convergence weights and bias parameters are continuously adjusted using actual labels to obtain the optimal weights and bias parameters for calculating the sound pickup state value, thereby improving the accuracy of the multilayer perceptron output. The actual labels are set by professionals in this field and will not be elaborated here.

[0025] The pickup status of each microphone is compared with the preset first stable state threshold and the preset second stable state threshold to determine the result. It should be noted that the preset first stable state threshold is greater than the preset second stable state threshold.

[0026] If the pickup state of each microphone is greater than the preset second stable state threshold and less than the preset first stable state threshold, then the microphone is determined to be a stable microphone. Otherwise, the microphone is determined to be an unstable microphone.

[0027] It should be explained that the preset first stable state threshold and the preset second stable state threshold are important parameters for determining whether the microphone is in a stable state. The preset second stable state threshold should be located in the area where the sensitivity begins to decrease or changes significantly, and can be set by offset amplitude. The preset first stable state threshold is usually set in the area where the microphone exhibits excessive sensitivity, that is, when the sensitivity changes too much, the microphone is prone to noise or over-response, and can be set by frequency response.

[0028] In step S3, the speech signal segment covered by the stable microphone is divided into multiple time windows, and a short-time Fourier transform is performed on each window to obtain a two-dimensional time-frequency diagram. The spectral edge information of speech signal segments is obtained by processing the two-dimensional time-frequency graph using a time-frequency domain convolution operator. The speech signal segments covered by the stable microphone are divided into low-frequency band, mid-frequency band and high-frequency band by a preset frequency band range. If the spectral edge information of a speech signal segment is less than the preset frequency band range, the speech signal segment covered by the stable microphone is classified as a low-frequency band; if the spectral edge information of a speech signal segment is within the preset frequency band range, the speech signal segment covered by the stable microphone is classified as a mid-frequency band; if the spectral edge information of a speech signal segment is greater than the preset frequency band range, the speech signal segment covered by the stable microphone is classified as a high-frequency band.

[0029] The power after noise reduction for each frequency band is obtained by spectral subtraction; The power spectrum of each frequency band is obtained by squaring each point in the two-dimensional time-frequency graph. The total power of each frequency band is obtained by summing the power of each time window in each frequency band. The noise power percentage of each frequency band is obtained by subtracting 1 from the ratio of the noise-reduced power of each frequency band to the total power of each frequency band and taking the absolute value. The energy intensity of the speech signal in each frequency band is obtained by summing the points of the power spectrum. It needs to be explained that the Short-Time Fourier Transform (SFT) is a method for analyzing the time-frequency characteristics of a signal. It extends the traditional Fourier Transform, simultaneously displaying the changes of a signal in time and frequency, and is used to obtain a two-dimensional time-frequency plot. A two-dimensional time-frequency plot is a two-dimensional image that can simultaneously show the changes of a signal in time and frequency. Each point on the horizontal axis of a two-dimensional time-frequency plot represents a different time window, and each point on the vertical axis represents the frequency of the signal. The time-frequency domain convolution operator is a tool for performing convolution operations on signals in the time-frequency domain. It combines the ideas of convolution operations and time-frequency analysis and is used to process and analyze non-stationary signals. Spectral edge information refers to information in the signal spectrum used to describe the frequency characteristics of the signal. It is usually used to describe the edge parts of the energy distribution in the spectrum, that is, the frequency boundaries of the energy concentration area, including the spectral edge frequencies. The preset frequency band range is an important parameter for determining the frequency band interval of a speech signal segment covered by a stable microphone. The preset frequency band range can be dynamically adjusted by monitoring the spectral characteristics of the signal and noise in real time. Spectral subtraction is a commonly used noise suppression technique. By subtracting the estimated noise spectrum from the signal spectrum, the signal quality is improved, and it is used to obtain the power after noise reduction for each divided frequency band. In step S4, the difference in the noise power ratio of adjacent time windows of each frequency band is taken as the noise change of each frequency band. The noise variation trend of each frequency band is obtained by averaging the noise variation of each frequency band. The noise variation trend and speech signal energy intensity of each frequency band are standardized to obtain the noise variation factor and speech intensity factor. It should be noted that the standardization methods include, but are not limited to, standard linear transformation based on interval scaling, statistical Z-Score standardization, or normalization based on nonlinear mapping functions. The application methods of standardization will not be elaborated here.

[0030] The noise reduction weighting index for each frequency band is calculated by combining the noise variation factor and the speech intensity factor. The calculation formula is as follows: ,in, For noise variation factor, For speech intensity factor, and To preset the weighting coefficients, This represents the noise reduction weighting index for each frequency band. It should be noted that the preset weighting coefficients are important parameters for balancing the influence of noise variation factors and speech intensity factors on the noise reduction weighting index of each frequency band. Verification is required through experimental data and test results in practical applications. Tests under different noise conditions and speech signals are needed to evaluate the accuracy of the noise reduction weighting index of each frequency band under different weighting coefficients, ensuring that the final weighting coefficients provide the optimal calculation effect. A larger noise variation factor and a stronger speech intensity factor result in more drastic fluctuations or intensity changes in noise within that frequency band, leading to a larger noise reduction weighting index for each frequency band. Conversely, a smaller noise variation factor and a weaker speech intensity factor result in smaller noise variations within that frequency band, leading to a smaller noise reduction weighting index for each frequency band.

[0031] The determination is based on the noise reduction weighting index of the frequency band division and the preset noise reduction threshold: If the noise reduction weight index of the frequency band division is greater than or equal to the preset noise reduction threshold, then the frequency band division is marked. If the noise reduction weight index of the frequency band division is less than the preset noise reduction threshold, it is determined that the frequency band division will not be marked.

[0032] It should be explained that the preset noise reduction threshold is an important parameter for determining whether to mark the frequency band. By testing the noise and speech signal characteristics under different environments, an initial threshold range is selected. Through experimental data analysis, the threshold is adjusted to find a suitable balance point, verify the noise reduction effect and speech quality of the system under different thresholds, ensure system stability and speech clarity, and further optimize the threshold setting based on actual tests and system feedback.

[0033] Finally, it should be noted that in this paper, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0034] Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0035] In this document, the singular forms “a,” “an,” and “the” may also include the plural forms unless the context clearly indicates otherwise. It should also be understood that terms such as “comprising / including” or “having” specify the presence of the stated features, integrals, steps, operations, components, parts, or combinations thereof, but do not preclude the possibility of the presence or addition of one or more other features, integrals, steps, operations, components, parts, or combinations thereof. Meanwhile, the term “and / or” as used in this specification includes any and all combinations of the associated listed items.

[0036] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0037] The above description of the disclosed embodiments will enable those skilled in the art to make or use various modifications to these embodiments. It will be readily apparent to those skilled in the art that the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech input noise reduction method based on deep learning, characterized in that: Includes the following steps: Step S1: Monitor the microphones used for voice acquisition, obtain the pointing angle and sensitivity change of each microphone, and calculate the pickup angle offset of each microphone based on the pointing angle. Step S2: Evaluate the pickup status of each microphone by combining the pickup angle offset amplitude and sensitivity change, filter the microphones according to the pickup status to obtain stable microphones, and retrieve the speech signal segment covered by the stable microphones. Step S3: Detect the spectral edge information of the speech signal segment using the time-frequency domain convolution operator, divide the speech signal segment into frequency bands based on the spectral edge information, and obtain the noise power ratio and speech signal energy intensity of each divided frequency band; Step S4: Evaluate the noise change trend of each divided frequency band based on the noise power ratio, generate the noise reduction weight index of each divided frequency band in combination with the speech signal energy intensity, and mark the divided frequency bands based on the noise reduction weight index.

2. The speech input noise reduction method based on deep learning according to claim 1, characterized in that: In step S1, the pointing angle of each microphone is obtained by a rotary encoder installed on the shaft of each microphone; The received sound wave signal is converted into an electrical signal by the built-in capacitive sensor of each microphone, which serves as the input signal for each microphone. The external sound wave signal is converted into an electrical signal by the dynamic sensor built into each microphone, which serves as the response signal for each microphone. The ratio of the response signal of each microphone to the input signal is used as the sensitivity of each microphone.

3. The speech input noise reduction method based on deep learning according to claim 2, characterized in that: In step S1, the absolute value of the difference between the sensitivity of each microphone and the preset standard sensitivity is taken as the sensitivity change of each microphone. The absolute value of the difference between the pointing angle of each microphone and the preset reference angle is used as the pickup angle offset of each microphone.

4. The speech input noise reduction method based on deep learning according to claim 3, characterized in that: In step S2, multiple sets of pickup angle offset amplitude and sensitivity change amount are obtained, and the pickup angle offset amplitude dataset and sensitivity change amount dataset are integrated. A multilayer perceptron model was constructed by combining the datasets of pickup angle offset amplitude and sensitivity change to analyze the pickup status of each microphone.

5. The speech input noise reduction method based on deep learning according to claim 4, characterized in that: In step S2, the preset first stable state threshold is greater than the preset second stable state threshold; If the pickup state of each microphone is greater than the preset second stable state threshold and less than the preset first stable state threshold, then the microphone is determined to be a stable microphone. Otherwise, the microphone is determined to be an unstable microphone.

6. The speech input noise reduction method based on deep learning according to claim 1, characterized in that: In step S3, the speech signal segment covered by the stable microphone is divided into multiple time windows, and a short-time Fourier transform is performed on each window to obtain a two-dimensional time-frequency diagram. The spectral edge information of speech signal segments is obtained by processing the two-dimensional time-frequency graph using a time-frequency domain convolution operator.

7. The speech input noise reduction method based on deep learning according to claim 6, characterized in that: In step S3, if the spectral edge information of the speech signal segment is less than the preset frequency band range, the speech signal segment covered by the stable microphone is divided into the low frequency band. If the spectral edge information of a speech signal segment is within the preset frequency band, then the speech signal segment covered by the stable microphone will be divided into the mid-frequency band. If the spectral edge information of a speech signal segment is greater than the preset frequency band range, then the speech signal segment covered by the stable microphone will be divided into a high-frequency band.

8. The speech input noise reduction method based on deep learning according to claim 7, characterized in that: In step S3, the noise-reduced power of each frequency band is obtained by spectral subtraction; The power spectrum of each frequency band is obtained by squaring each point in the two-dimensional time-frequency graph. The total power of each frequency band is obtained by summing the power of each time window in each frequency band. The noise power percentage of each frequency band is obtained by subtracting 1 from the ratio of the noise-reduced power of each frequency band to the total power of each frequency band and taking the absolute value. The energy intensity of the speech signal in each frequency band is obtained by summing the points of the power spectrum.

9. The speech input noise reduction method based on deep learning according to claim 8, characterized in that: In step S4, the difference in the noise power ratio of adjacent time windows of each frequency band is taken as the noise change of each frequency band. The noise variation trend of each frequency band is obtained by averaging the noise variation of each frequency band. The noise variation trend and speech signal energy intensity of each frequency band are standardized to obtain the noise variation factor and speech intensity factor. The noise reduction weight index for each frequency band is calculated by combining the noise variation factor and the speech intensity factor.

10. The speech input noise reduction method based on deep learning according to claim 9, characterized in that: In step S4, if the noise reduction weight index of the divided frequency band is greater than or equal to the preset noise reduction threshold, it is determined that the divided frequency band will be marked; if the noise reduction weight index of the divided frequency band is less than the preset noise reduction threshold, it is determined that the divided frequency band will not be marked.