An intelligent sound bar voice denoising method for a construction area

By separating voice and noise acquisition modes in a smart speaker and combining deep learning and adaptive filter technology, the problem of voice noise reduction in complex noise environments in construction areas has been solved, achieving efficient and universal voice noise reduction effect.

CN116564328BActive Publication Date: 2026-02-13ZHEJIANG DAYOU INDUSTRIAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310353824.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-04
Publication Date
2026-02-13
Estimated Expiration
2043-04-04

AI Technical Summary

Technical Problem

In construction areas, existing technologies struggle to effectively distinguish between noise and speech, resulting in unsatisfactory speech noise reduction. Furthermore, different noise reduction technologies are only applicable to specific noise environments, hindering their widespread adoption.

Method used

The smart speaker is divided into voice acquisition mode and noise acquisition mode. Voice preprocessing and secondary noise reduction are performed through spectrum extraction and deep learning models, and noise cancellation is performed by combining adaptive filters and gated loop units.

Benefits of technology

It achieves efficient voice noise reduction in construction areas, maintaining the naturalness and sound quality of speech, and is suitable for various noise types and environments, thus improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116564328B_ABST
    Figure CN116564328B_ABST
Patent Text Reader

Abstract

The application discloses a kind of intelligent sound box voice denoising methods for construction area, comprising: the acquisition mode is divided into voice acquisition mode and noise acquisition mode;Voice acquisition mode is closed, and noise acquisition mode is intermittently started, and the spectrum of noise is extracted;Voice acquisition mode is started, and voice sound source is collected, and according to the spectrum extraction result of noise, the voice is preprocessed;The voice after preprocessing is secondary denoising;The voice after denoising is played, transmitted or saved as the effective audio of intelligent sound box.The application divides the acquisition mode into voice acquisition mode and noise acquisition mode, can carry out noise acquisition and processing in the gap of voice acquisition, can reduce the separation work of noise and voice, more accurately obtain the relevant characteristics of noise, to facilitate subsequent processing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of noise reduction, in particular to a voice de-noising method for an intelligent sound box in a construction area. BACKGROUND

[0002] With the development of voice technology, voice recognition is gradually popularized in daily life. However, in various scenes in daily use, due to the existence of various noises and the interference of device signals, the voice quality and intelligibility will be affected, and the performance of the voice recognition system will be sharply reduced.

[0003] In the prior art, the technical solutions for de-noising voice mainly include: a method based on spectral subtraction, which converts the voice signal from the time domain to the frequency domain, and then removes the influence of the noise signal from the frequency spectrum signal; a method based on filter, which reduces the influence of noise signal by designing a special de-noising filter.

[0004] However, in the construction area, background noise is always present, so the above-mentioned solutions have the following shortcomings: it is difficult to estimate the noise in the noisy voice, and the effect is not ideal; the existing technology is easy to cause information loss and distortion of the voice signal when operating on the frequency spectrum, which affects the intelligibility and naturalness of the voice; different voice de-noising technologies are only suitable for specific noise environment and type, and the technology popularization is poor. SUMMARY

[0005] In view of the problem that the prior art is difficult to distinguish noise and voice in the construction area, the present application provides a voice de-noising method for an intelligent sound box in a construction area, which separates noise and voice with noise by voice collection mode and noise collection mode, reduces the difficulty of noise processing, and provides a more effective de-noising way for noisy voice.

[0006] The technical solution of the present application is as follows.

[0007] A voice de-noising method for an intelligent sound box in a construction area, comprising:

[0008] dividing the collection mode into voice collection mode and noise collection mode;

[0009] when the voice collection mode is off, the noise collection mode is intermittently started to extract the spectrum of noise;

[0010] when the voice collection mode is started, the voice sound source is collected, and the voice is preprocessed according to the spectrum extraction result of the noise; the preprocessed voice is de-noised again;

[0011] the de-noised voice is played, transmitted or saved as effective audio of the intelligent sound box.

[0012] Because the noise composition of the construction area is relatively complex, sometimes it is difficult to distinguish noise and voice, because the voice of many workers can also be a kind of noise, the application divides the collection mode into voice collection mode and noise collection mode, can collect and process noise in the gap of voice collection, can reduce the separation work of noise and voice, more accurately obtain the related characteristics of noise, in order to facilitate subsequent processing.

[0013] As preferred, the voice collection mode is closed, and the noise collection mode is intermittently started, comprising:

[0014] The preset noise collection condition is used to determine whether the noise collection condition is met in real time when the voice collection mode is closed, and if so, the noise collection mode is started to collect the noise of the construction area.

[0015] As preferred, the noise is subjected to frequency spectrum extraction, comprising:

[0016] The noise is grouped and subjected to Fourier transform, and the obtained group is averaged to obtain the frequency spectrum of the noise.

[0017] The noise data is subjected to short-time Fourier transform to obtain spectrum graph information, and the spectrum graph information comprises the amplitude and phase of the noise data.

[0018] The basic idea of STFT transform is to truncate the non-stationary signal with a window function, regard the signal in the truncated window as stationary, and then perform Fourier transform on the series of short-time stationary signals to obtain a two-dimensional time-frequency matrix.

[0019] The signal x(t) belongs to L 2 (R), and its mathematical expression is:

[0020]

[0021] In the formula, A is the amplitude, ω is the angular frequency, and ψ is the initial phase angle. The STFT transform of the signal x(t) is defined as follows:

[0022]

[0023] Where g(t-τ) is the window function, and the window function slides on the time axis with the continuous change of τ to analyze the signal. Once the window function is determined in the short-time Fourier transform, the entire time-frequency window remains unchanged, that is, it has only a single resolution, and the short-time Fourier transform spectrum analysis result is affected by the signal interception position and the window length.

[0024] As preferred, the noise collection condition comprises:

[0025] The current decibel value is obtained, and when the current decibel value is greater than a first preset value, a first condition is triggered and lasts for T time.

[0026] acquiring a decibel difference value in a preset time interval, triggering a second condition when the decibel difference value is greater than a second preset value, and triggering a third condition when the decibel difference value is less than or equal to the second preset value, wherein T is greater than t;

[0027] determining whether a current time is in a triggering state, such as triggering the first condition and the second condition, starting the noise collection mode in a first period, such as triggering the first condition and the third condition, starting the noise collection mode in a second period, and starting the noise collection mode in the second period when the first condition is not triggered, wherein the first period is less than the second period.

[0028] When the absolute value of the decibel is large, the noise is strong, and when the difference value of the decibel is large, the noise is unstable, so according to the characteristics of the noise, the noise is collected, which can be more targeted and reduce the consumption of invalid energy and improve the utilization efficiency.

[0029] As a preferred, the pre-processing of the voice according to the frequency spectrum extraction result of the noise includes:

[0030] The frequency spectrum extraction result of the noise is used to pre-process the voice, and the pre-processed voice is obtained.

[0031] As a preferred, the pre-processing of the voice according to the frequency spectrum extraction result of the noise includes:

[0032] The pre-processed voice is format processed to obtain a voice of a predetermined format, then the voice of the predetermined format is sampled according to a sampling rate in the predetermined format to obtain sampling point information, the deep learning noise reduction model is used to process the obtained sampling point information to generate noise-reduced sampling point information, and finally the voice after secondary noise reduction is generated according to the noise-reduced sampling point information.

[0033] The core operation mode of deep learning is:

[0034]

[0035] The result of the convolution layer output enters the down-sampling layer, and each feature output by the convolution layer is down-sampled. That is, each feature is respectively subjected to weighted summation operation or maximum value operation, multiplied by a multiplier bias, added to an additional bias, and then subjected to an excitation function, and finally a feature map is obtained in the down-sampling layer. The calculation formula is:

[0036]

[0037] Wherein: down() is a down-sampling function, β is a multiplier bias of down-sampling, and b is a corresponding additional bias. Then, the calculation of the repeated convolution layer and the down-sampling layer is repeated, which is a C-S process. Finally, a vector is output from the network.

[0038] The training process of the deep learning denoising model comprises:

[0039] Obtain sample voice data, wherein the sample voice data is voice data obtained by mixing noise-free voice data and known noise data; extract sample spectrum graph information of the sample voice data, and calculate a signal-to-noise ratio corresponding to the sample spectrum graph information; input the sample spectrum graph information into a preset model, train the preset model until a signal-to-noise ratio output by the preset model is the signal-to-noise ratio corresponding to the sample spectrum graph information, and determine the trained preset model as the neural network model.

[0040] The substantial effects of the present application include:

[0041] Using the sampling point information as the input and output of the deep learning denoising model does not require complex operations such as noise estimation, is simple to implement, does not cause distortion problems such as voice noise, has better naturalness and sound quality, and brings better user experience; in addition, the deep learning denoising model is suitable for various noise types and environments through learning of a large amount of noisy voice and clean voice, has universal applicability, and is convenient for promotion. BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1 is a flowchart of an embodiment of the present application. DETAILED DESCRIPTION

[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will combine embodiments to clearly and completely describe the technical solutions of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0044] It should be understood that in various embodiments of the present application, the magnitude of the serial number of each process does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0045] It should be understood that in the present application, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.

[0046] It should be understood that in the present application, "a plurality of" means two or more. "And / or" is only a description of the association between the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally represents that the associated objects before and after are in an "or" relationship. "Including A, B and C", "including A, B, C" means that A, B and C are all included, "including A, B or C" means that one of A, B and C is included, and "including A, B and / or C" means that any one or any two or three of A, B and C is included.

[0047] The technical solutions of the present application will be described in detail below with specific examples. The examples can be combined with each other, and the same or similar concepts or processes can not be described in some examples.

[0048] Embodiment:

[0049] A method for intelligent sound box voice denoising in a construction area, as shown in Figure 1 , comprising:

[0050] The collection mode is divided into a voice collection mode and a noise collection mode;

[0051] When the voice collection mode is off, the noise collection mode is intermittently started, and the noise is spectrum extracted;

[0052] When the voice collection mode is started, the voice sound source is collected, and the voice is preprocessed according to the spectrum extraction result of the noise; the preprocessed voice is denoised again;

[0053] The denoised voice is played, transmitted or saved as an effective audio of the intelligent sound box.

[0054] In this embodiment, the collection mode is divided into a voice collection mode and a noise collection mode, noise collection and processing can be performed in the gap of voice collection, which can reduce the separation work of noise and voice, and more accurately obtain the relevant characteristics of noise for subsequent processing.

[0055] In this embodiment, when the voice collection mode is off, the noise collection mode is intermittently started, comprising:

[0056] A preset noise collection condition is set, and when the voice collection mode is off, it is judged in real time whether the noise collection condition is met, and if so, the noise collection mode is started to collect the noise in the construction area.

[0057] In this embodiment, the spectrum extraction of the noise comprises:

[0058] The noise is grouped and Fourier transformed, and the obtained group is averaged to obtain the spectrum of the noise;

[0059] The noise data is subjected to a short-time Fourier transform (STFT) to obtain spectral graph information including the amplitude and phase of the noise data.

[0060] The basic idea of the STFT transform is to truncate a non-stationary signal with a window function, to regard the truncated signal in the window as stationary, and to perform Fourier transform on the series of short-time stationary signals to obtain a two-dimensional time-frequency matrix.

[0061] The signal x(t) is in L 2 (R), and the mathematical expression is:

[0062]

[0063] In the formula, A is the amplitude, ω is the angular frequency, and ψ is the initial phase angle. The STFT transform of the signal x(t) is defined as follows:

[0064]

[0065] In the formula, g(t-τ) is a window function, and the window function slides on the time axis with the continuous change of τ to analyze the signal. Once the window function is determined in the short-time Fourier transform, the entire time-frequency window remains unchanged, i.e., it has only a single resolution, and the short-time Fourier transform spectrum analysis result is affected by the signal truncation position and the window length.

[0066] In this embodiment, the noise collection condition includes:

[0067] The current decibel value is obtained, and when the current decibel value is greater than a first preset value, a first condition is triggered and lasts for a time T.

[0068] A decibel difference value of a preset time interval is obtained, and when the decibel difference value is greater than a second preset value, a second condition is triggered and lasts for a time t, and when the decibel difference value is less than or equal to the second preset value, a third condition is triggered and lasts for a time T, where T is greater than t.

[0069] It is determined whether the current moment is in a triggered state, such as the first condition and the second condition being triggered, and the noise collection mode is started in a first period, such as the first condition and the third condition being triggered, and the noise collection mode is started in a second period, and if the first condition is not triggered, the noise collection mode is started in the second period, where the first period is less than the second period.

[0070] The absolute value of the decibel is large, and the noise is strong, and the difference value of the decibel is large, and the noise is unstable, so according to the characteristics of the noise, the noise is collected, which can be more targeted, and the energy consumption is reduced, and the utilization efficiency is improved.

[0071] In this embodiment, the pre-processing of the voice according to the spectral extraction result of the noise includes:

[0072] The speech is inversely compensated using the spectrum extraction result of the noise to obtain the speech after noise reduction.

[0073] In this embodiment, the secondary noise reduction of the preprocessed speech comprises:

[0074] The preprocessed speech is format-processed to obtain speech in a predetermined format (for example, the PCM format of 16,000 Hz sampling rate, 16-bit quantization and single channel), and then the speech in the predetermined format is sampled at the sampling rate in the predetermined format to obtain sampling point information, the obtained sampling point information is processed by the deep learning noise reduction model to generate noise-reduced sampling point information, and finally the speech after secondary noise reduction is generated according to the noise-reduced sampling point information.

[0075] The deep learning noise reduction model is trained in advance through learning of a large amount of noisy speech and clean speech, and the training process of the deep learning noise reduction model comprises:

[0076] Obtain sample speech data, wherein the sample speech data is speech data obtained by mixing noise-free speech data and known noise data; extract sample spectrum graph information of the sample speech data, and calculate the signal-to-noise ratio corresponding to the sample spectrum graph information; input the sample spectrum graph information into a preset model, train the preset model until the signal-to-noise ratio output from the preset model is the signal-to-noise ratio corresponding to the sample spectrum graph information, and then determine the trained preset model as the deep learning noise reduction model.

[0077] The obtained sampling point information is processed, and the processing specifically comprises:

[0078] The core operation mode of deep learning is:

[0079]

[0080] The result output by the convolution layer enters the down-sampling layer, and each feature output by the convolution layer is down-sampled. That is, each feature is respectively subjected to weighted summation operation or maximum value operation, multiplied by a multiplier bias, added with an additional bias, and then subjected to an excitation function, and finally a feature map is obtained in the down-sampling layer. The calculation formula is:

[0081]

[0082] Wherein, down() is a down-sampling function, β is a multiplier bias of down-sampling, and b is a corresponding additional bias. Then, the calculation of the repeated convolution layer and the down-sampling layer is repeated, which is a C-S process. Finally, a vector is output from the network. In this embodiment, the smart speaker is mainly used to provide an amplification effect for a commander in a construction area.

[0083] In this embodiment, the noise estimation is performed using the method of minimum control regression average (MCRA). MCRA is a method combining regression average and minimum value tracking. First, the probability of speech presence is determined using local minimum value. Then, according to the probability of speech presence, it is determined which frequency band is used for noise estimation, and the noise estimation is obtained using the regression average method.

[0084] In addition, an adaptive filter is used to automatically adjust the performance of digital signal processing according to the input signal. In contrast, a non-adaptive filter has static filter coefficients, which together form a transfer function.

[0085] For some applications, it is required to use adaptive coefficients for processing because the parameters required for operation are not known in advance, such as the characteristics of some noise signals. In this case, an adaptive filter is usually used, which uses feedback to adjust the filter coefficients and the frequency response. In general, the adaptive process involves an algorithm that uses a cost function to determine how to change the filter coefficients to reduce the cost of the next iteration. The value function is a criterion for the best performance of the filter, such as the ability to reduce the noise component in the input signal. With the increasing performance of digital signal processors, adaptive filters are becoming more and more common, and today they are widely used in mobile phones and other communication devices.

[0086] The principle of noise cancellation of an adaptive filter is to optimally estimate the noise by minimizing the mean square error or variance, and then subtract the optimally estimated noise from the speech containing noise to achieve the purpose of noise reduction, improving the signal-to-noise ratio, and enhancing the speech.

[0087] Therefore, the combination of speech signal processing and deep learning can automatically learn rules from a large amount of data, and the more noise heard, the better the speech processing, which in turn can help the speech noise reduction processing to estimate parameters.

[0088] The structure processing module is mainly implemented by a good performance and low resource demand gated recurrent unit (GRU), which is responsible for data storage and network calculation. This functional architecture is aimed at construction area noise cancellation, which consists of three functional modules: speech detection, noise detection, and noise cancellation. Speech detection detects speech signals in real time, and only when speech signals are detected will noise suppression be performed. At this time, noise filtering and cancellation are performed on the signals after adaptive filtering. Noise cancellation, most noise signals have a wide bandwidth and smooth spectrum, so the speech signal is compared with the noise signal for filtering, and the noise part in the sound signal is eliminated. In addition, other sound artifacts are also avoided as much as possible.

[0089] The filtered audio signal is used as the input and output of the deep learning model, which rarely causes distortion of the voice, has better naturalness and sound quality, and brings better user experience.

[0090] In addition, the adaptive filter combined with the deep learning neural network scheme in the present application considers adding a filter in advance compared with the traditional deep learning neural network sound processing scheme, does not need to perform noise estimation, directly compares and eliminates the processed voice and the processed noise, consumes relatively less resources while having good effects, learns a large number of noises and clean voices in the construction process, has universal applicability for the current construction scene, and is convenient for promotion.

[0091] Through the description of the above embodiments, those skilled in the art can understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the specific device is divided into different functional modules to complete all or part of the functions described above.

[0092] In the embodiments provided in the present application, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the above-described embodiments of the structure are only illustrative, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another structure, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, structures or units, which can be electrical, mechanical or other forms.

[0093] The units described as separate components can or can not be physically separated, and the components shown as units can be one physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.

[0094] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0095] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The software product is stored in a storage medium, including a plurality of instructions to make a device (which can be a single-chip microcomputer, a chip, etc.) or a processor execute all or part of the steps of the various embodiments of the method of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0096] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for voice noise reduction of smart speakers used in construction areas, characterized in that, include: The acquisition modes are divided into voice acquisition mode and noise acquisition mode; When the voice acquisition mode is off, the noise acquisition mode is intermittently activated to extract the spectrum of the noise. When the voice acquisition mode is activated, the voice source is acquired, and the voice is preprocessed based on the noise spectrum extraction results. Secondary noise reduction is performed on the preprocessed speech; The noise-reduced speech is played, transmitted, or saved as valid audio by the smart speaker; When the voice acquisition mode is turned off, the noise acquisition mode is intermittently activated, including: preset noise acquisition conditions, and when the voice acquisition mode is turned off, it is determined in real time whether the noise acquisition conditions are met. If they are met, the noise acquisition mode is activated to collect noise from the construction area. The noise acquisition conditions include: acquiring the current decibel value; when the current decibel value is greater than a first preset value, triggering a first condition for a duration of T; acquiring the decibel difference over a preset time interval; when the decibel difference is greater than a second preset value, triggering a second condition for a duration of t; when the decibel difference is less than or equal to the second preset value, triggering a third condition for a duration of T, where T is greater than t; determining whether the current condition is in a triggered state; if the first and second conditions have been triggered, starting the noise acquisition mode with a first cycle; if both the first and third conditions have been triggered, starting the noise acquisition mode with a second cycle; if the first condition has not been triggered, starting the noise acquisition mode with a second cycle, where the first cycle is less than the second cycle. The secondary noise reduction of the preprocessed speech includes: format processing the preprocessed speech to obtain speech in a predetermined format; sampling the speech in the predetermined format according to the sampling rate in the predetermined format to obtain sampling point information; performing noise reduction processing on the obtained sampling point information through a deep learning noise reduction model to generate noise-reduced sampling point information; and finally generating secondary noise-reduced speech based on the noise-reduced sampling point information.

2. The method for voice noise reduction of a smart speaker in a construction area according to claim 1, characterized in that, The noise spectrum extraction includes: The noise is grouped, subjected to Fourier transform, and the average of the resulting groups is calculated to obtain the noise spectrum. The noise data is subjected to a short-time Fourier transform to obtain a spectrum information, which includes the amplitude and phase of the noise data.

3. The method for voice noise reduction of a smart speaker in a construction area according to claim 1, characterized in that, The preprocessing of the speech based on the noise spectrum extraction results includes: By using the noise spectrum extraction results, the speech is reverse-compensated to obtain the noise-reduced speech.

Citation Information

Patent Citations

  • Voice noise reduction algorithm

    CN108428456A

  • Voice noise reduction method and device, equipment and storage medium

    CN114242104A