Speech noise reduction methods, systems, devices and computer storage media
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]本申请的主要目的在于提供一种语音降噪方法、系统、设备及计算机存储介质,旨在解决语音采集的效果不佳的技术问题
[0016]本申请实施例提供了一种语音降噪方法,应用于语音降噪系统,语音降噪系统包括第一采集模组和第二采集模组,通过获取第一采集模组的第一采集数据和第二采集模组的第二采集数据,根据第一采集数据和第二采集数据确定目标语音数据,这种语音降噪方法同时对第一采集模组采集的第一采集数据和第二采集模组采集的第二采集数据进行处理,以结合第一采集数据和第二采集数据依次确定目标语音数据,因为整个语音处理流程并非单一通道语音采集处理,即不会存在必须依赖对模组工作状态的采集需求,进而避免第一采集模组的检测准确性直接影响到最终语音采集效果的缺陷,也就是这种语音降噪方法通过同时结合第一采集数据和第二采集数据进行处理得到目标语音数据,进而可以大大降低对第一采集模组的检测的依赖,即无论第一采集模组处于何种状态,均会使用第一采集模组的第一采集数据进行协同处理,以保证语音采集的效果。
Smart Images

Figure CN122575394A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech noise reduction technology, and in particular to a speech noise reduction method, system, device and computer storage medium. Background Technology
[0002] With the continuous development of voice acquisition terminals with two voice acquisition channels, users have also put forward higher requirements for the voice acquisition methods of voice acquisition terminals.
[0003] Traditional voice acquisition methods involve detecting the operational status of two acquisition channels. These channels typically include a first acquisition module on the voice acquisition terminal and a second acquisition module (located differently from the first) that communicates with the terminal. When the first acquisition module is detected as active, only it is used for voice pickup and noise reduction. When the first acquisition module is detected as idle, the system switches to the second acquisition module for acquisition and noise reduction. This method has limitations. The accuracy of the first acquisition module's detection directly affects the final voice acquisition result (e.g., erroneous information from the detection sensor may lead to a misjudgment of the first acquisition module's status). In other words, the accuracy of the first acquisition module's detection directly impacts the final voice acquisition result, leading to poor overall voice acquisition quality.
[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide a speech noise reduction method, system, device and computer storage medium, which aims to solve the technical problem of poor speech acquisition effect.
[0006] To achieve the above objectives, this application provides a speech denoising method, which is applied to a speech denoising system. The speech denoising system includes a first acquisition module and a second acquisition module. The speech denoising method includes: Acquire the first acquisition data from the first acquisition module and the second acquisition data from the second acquisition module, and determine the target speech data based on the first acquisition data and the second acquisition data.
[0007] In one embodiment, the step of determining the target speech data based on the first acquired data and the second acquired data includes: The current sound energy data for the current time frame is determined based on the first and second collected data. When the current sound energy data does not match the preset user voice energy, noise energy data is determined based on the current sound energy data and the historical sound energy data of the previous frame. The signal-to-noise ratio (SNR) weight of the current time frame is determined based on the noise energy data and the current sound energy data, and the target speech data is determined based on the SNR weight, the first acquired data, and the second acquired data.
[0008] In one embodiment, the current sound energy data includes first sound energy data of the current time frame in the first acquisition data and second sound energy data of the current time frame in the second acquisition data. The step of determining the current sound energy data of the current time frame based on the first acquisition data and the second acquisition data includes: Determine the first current frame data of the current time frame in the first collected data, and determine the first overall energy data of the current time frame based on the first current frame data. Use the quotient of the first overall energy data and the number of sampling points of the current time frame as the first sound energy data. The second current frame data of the current time frame in the second acquisition data is determined, and the second overall energy data of the current time frame is determined based on the second current frame data. The quotient of the second overall energy data and the number of sampling points of the current time frame is taken as the second sound energy data.
[0009] In one embodiment, the noise energy data includes first noise energy data of the current time frame in the first acquired data and second noise energy data of the current time frame in the second acquired data. The step of determining the noise energy data based on the current sound energy data and the historical sound energy data of the previous frame includes: The first historical sound energy data of the first collected data in the historical sound energy data of the previous frame is determined, the first product between the first historical sound energy data and the preset first factor is determined, the second product between the first sound energy data and the preset second factor in the current sound energy data is determined, and the sum of the first product and the second product is taken as the first noise energy data, wherein the sum of the first factor and the second factor is one; The second historical sound energy data of the second acquired data in the historical sound energy data of the previous frame is determined, the third product between the second historical sound energy data and the preset first factor is determined, the fourth product between the second sound energy data in the current sound energy data and the preset second factor is determined, and the sum of the third product and the fourth product is taken as the second noise energy data.
[0010] In one embodiment, the step of determining the signal-to-noise ratio weight of the current time frame based on the noise energy data and the current sound energy data includes: A first signal-to-noise ratio is determined based on the first noise energy data in the noise energy data and the first sound energy data in the current sound energy data, and a first power ratio is determined based on the first signal-to-noise ratio; A second signal-to-noise ratio (SNR) is determined based on the second noise energy data in the noise energy data and the second sound energy data in the current sound energy data, and a second power ratio is determined based on the second SNR. The SNR weight of the current time frame is determined based on the second power ratio and the first power ratio.
[0011] In one embodiment, the signal-to-noise ratio (SNR) weight includes a first SNR weight of the first acquired data and a second SNR weight of the second acquired data. The step of determining the SNR weight of the current time frame based on the second power ratio and the first power ratio includes: The first signal-to-noise ratio weight is determined based on the second power ratio, the first power ratio, and a preset time constant, and the second signal-to-noise ratio weight corresponding to the first signal-to-noise ratio weight is determined, wherein the sum of the second signal-to-noise ratio weight and the first signal-to-noise ratio weight is one.
[0012] In one embodiment, the step of determining the target speech data based on the signal-to-noise ratio weights, the first acquired data, and the second acquired data includes: A first target speech signal is determined based on the first signal-to-noise ratio weight in the signal-to-noise ratio weight and the first acquired data, and a second target speech signal is determined based on the second signal-to-noise ratio weight in the signal-to-noise ratio weight and the second acquired data; Target speech data is determined based on the first target speech signal and the second target speech signal.
[0013] Furthermore, to achieve the above objectives, this application also provides a speech noise reduction system, which includes a controller, a first acquisition module, and a second acquisition module. The controller is connected to the first acquisition module and the second acquisition module, and the controller includes: The speech processing module is used to acquire the first acquisition data of the first acquisition module and the second acquisition data of the second acquisition module, and determine the target speech data based on the first acquisition data and the second acquisition data.
[0014] In addition, to achieve the above objectives, this application also provides a speech noise reduction device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the speech noise reduction method described above.
[0015] In addition, to achieve the above objectives, this application also provides a computer storage medium storing a speech denoising program, wherein when the speech denoising program is executed by a processor, it implements the steps of the speech denoising method described above.
[0016] This application provides a speech denoising method applied to a speech denoising system. The speech denoising system includes a first acquisition module and a second acquisition module. By acquiring first acquisition data from the first acquisition module and second acquisition data from the second acquisition module, target speech data is determined based on the first and second acquisition data. This speech denoising method simultaneously processes the first acquisition data acquired by the first acquisition module and the second acquisition data acquired by the second acquisition module to sequentially determine the target speech data by combining the first and second acquisition data. Because the entire speech processing flow is not a single-channel speech acquisition process, there is no need to rely on the acquisition requirements of the module's working state. This avoids the defect that the detection accuracy of the first acquisition module directly affects the final speech acquisition effect. In other words, this speech denoising method obtains the target speech data by simultaneously combining the first and second acquisition data, thereby greatly reducing the dependence on the detection of the first acquisition module. That is, regardless of the state of the first acquisition module, the first acquisition data of the first acquisition module will be used for collaborative processing to ensure the effect of speech acquisition. Attached Figure Description
[0017] Figure 1 This is a schematic flowchart of the first embodiment of the speech noise reduction method of this application; Figure 2 This is a schematic diagram of the speech noise reduction system of this application; Figure 3 This is a schematic flowchart of the second embodiment of the speech noise reduction method of this application; Figure 4 This is a schematic diagram of the overall process of the speech noise reduction method of this application; Figure 5 This is a schematic diagram of the controller module in this application; Figure 6 This is a schematic diagram of the speech noise reduction device in this application.
[0018] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings.
[0019] Explanation of icon numbers: 1001. Processing device; 1002. Read-only memory; 1003. Storage device; 1004. Random access memory; 1005. Bus; 1006. Input / output interface; 1007. Input device; 1008. Output device; 1009. Communication device; 1A-1N, Microphone A-Microphone N; 1a-1n, Microphone a-Microphone n; 10. First acquisition module; 20. Second acquisition module; 30. Analog-to-digital converter; 41. First synchronous serial port; 42. Second synchronous serial port; 50. Digital signal processor. Detailed Implementation
[0020] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0021] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0022] First, a dual-channel voice acquisition terminal refers to a voice acquisition terminal that uses two acquisition modules positioned at different locations to acquire sound, such as headphones or smart helmets. Taking headphones as an example, the two acquisition modules can be mounted on the earcups and microphone boom, or on the earcups and a handheld microphone, and there is communication between the two modules. Taking the acquisition module on the earcups and microphone boom as an example, the microphone boom can be bent or rotated to be adjusted to a suitable position for voice acquisition; it is also detachable, and can be installed on one side of the headphones via a USB (Universal Serial Bus) or 3.5mm audio jack interface, and can be removed. The acquisition module can be an array composed of microphones or other sound acquisition devices. It is worth noting that this application uses headphones as examples for illustration.
[0023] In headphones, uplink noise reduction typically uses a pickup module located on the earcups (i.e., the second pickup module) for sound pickup and processing. In noisy environments, a microphone boom (i.e., the first pickup module) is often needed to achieve better uplink noise reduction. The use of these two pickup modules is generally determined by detecting whether the microphone boom is plugged in or unplugged. However, this method suffers from inaccurate detection and speech stuttering caused by switching between the two pickup modules. Furthermore, if the pickup module is poorly positioned or the microphone boom's pickup module malfunctions, this method will still use the microphone boom's pickup module, thus affecting the overall speech pickup quality.
[0024] Therefore, based on the shortcomings of the above-mentioned speech acquisition methods, the speech denoising method of this application is proposed. The solution of this application embodiment is: by simultaneously processing the first acquisition data acquired by the first acquisition module and the second acquisition data acquired by the second acquisition module, the target speech data is determined sequentially by combining the first acquisition data and the second acquisition data. Because the entire speech processing flow is not a single-channel speech acquisition process, there is no need to rely on the acquisition requirements of the module's working state, thereby avoiding the defect that the detection accuracy of the first acquisition module directly affects the final speech acquisition effect. In other words, this speech denoising method obtains the target speech data by simultaneously combining the first acquisition data and the second acquisition data, thereby greatly reducing the dependence on the detection of the first acquisition module. That is, regardless of the state of the first acquisition module, the first acquisition data of the first acquisition module will be used for collaborative processing to ensure the effect of speech acquisition.
[0025] It should be noted that the executing entity in this embodiment can be a computing server with data processing, network communication, and program execution functions, such as a headset, head-mounted display, and smart helmet. The following description uses a headset as an example to illustrate this embodiment and the subsequent embodiments.
[0026] Based on this, embodiments of this application provide a speech noise reduction method, referring to... Figure 1 , Figure 1 This is a schematic flowchart of the first embodiment of the speech noise reduction method of this application.
[0027] Reference Figure 1 This application provides a speech denoising method, which is applied to a speech denoising system. The speech denoising system includes a first acquisition module and a second acquisition module. The speech denoising method includes: Step S10: Obtain the first acquisition data of the first acquisition module and the second acquisition data of the second acquisition module, and determine the target speech data based on the first acquisition data and the second acquisition data.
[0028] For example, the speech noise reduction method is applied to a speech noise reduction system, which includes a first acquisition module and a second acquisition module. Taking a pair of headphones as an example, see [reference needed]. Figure 2 , Figure 2This is a schematic diagram of the speech noise reduction system of this application. The first acquisition module 10 can be equipped with microphones including microphones A1A to N1N, which are then mounted on the microphone boom of the headset. The second acquisition module 20 can be equipped with microphones including microphones a1a to n1n, which are then mounted inside the earcups of the headset. This forms a first acquisition module and a second acquisition module positioned at different locations on the headset. The data acquired by the first and second acquisition modules are processed by the ADC (Analog-to-Digital Converter) 30 and then input to the DSP (Digital Signal Processor) 50 via the first synchronous serial port 41 and the second synchronous serial port 42, respectively, for speech signal processing, thereby completing the transmission and processing of the two acquired signals. It is worth noting that the ADC input needs to be pulled up to maintain a high-level output when there is no load. Regardless of whether the microphone boom is installed or removed, both pickup logic links exist. If the microphone boom is installed, the speech noise reduction processing method of this application continues to be executed; otherwise, the first acquisition data collected by the first acquisition module inside the earcup is directly processed and output, i.e., the first acquisition data is processed separately. At this time, by processing the two acquisition data simultaneously, the problem of poor speech acquisition effect caused by inaccurate microphone detection is reduced.
[0029] In this embodiment, the target speech data is obtained by simultaneously acquiring first acquisition data from the first acquisition module and second acquisition data from the second acquisition module, and then processing these two data sets. The first acquisition data refers to the data signal acquired by the first acquisition module, such as a continuous time-domain signal y. boom (t), the second acquired data refers to the data signal acquired by the second acquisition module, such as the continuous time domain signal y. ear (t), and then a series of processes are performed on these two continuous time-domain signals to obtain the target speech data, which refers to the speech data that needs to be uploaded in the end. At this time, the collected data from the two locations can be processed simultaneously to avoid the defect that the detection accuracy of the first acquisition module directly affects the final speech acquisition effect, thereby improving the effect of speech acquisition and uploading.
[0030] In this embodiment, a speech denoising method is provided, applied to a speech denoising system. The speech denoising system includes a first acquisition module and a second acquisition module. By acquiring first acquisition data from the first acquisition module and second acquisition data from the second acquisition module, target speech data is determined based on the first and second acquisition data. This speech denoising method simultaneously processes the first acquisition data acquired by the first acquisition module and the second acquisition data acquired by the second acquisition module to sequentially determine the target speech data by combining the first and second acquisition data. Because the entire speech processing flow is not a single-channel speech acquisition process, there is no requirement to rely on the acquisition of the module's working state. This avoids the defect that the detection accuracy of the first acquisition module directly affects the final speech acquisition effect. In other words, this speech denoising method obtains the target speech data by simultaneously combining the first and second acquisition data, thereby greatly reducing the dependence on the detection of the first acquisition module. That is, regardless of the state of the first acquisition module, the first acquisition data of the first acquisition module will be used for collaborative processing to ensure the effect of speech acquisition.
[0031] Furthermore, based on the first embodiment of this application described above, a second embodiment of the speech noise reduction method of this application is proposed. In this embodiment, reference is made to... Figure 3 , Figure 3 This is a flowchart illustrating a second embodiment of the speech noise reduction method of this application. The step of determining the target speech data based on the first and second acquired data includes: Step S11: Determine the current sound energy data of the current time frame based on the first and second acquisition data; In this embodiment, first acquisition data from the first acquisition module and second acquisition data from the second acquisition module are acquired. Simultaneously, the first and second acquisition data are processed into frames to obtain acquisition data for each frame. Then, the sound energy data of the current time frame can be determined based on the frame-segmented acquisition data. The frame-segmentation process can be achieved using a Hanning window to reduce spectral leakage, or other methods. In this case, the two acquisition data are processed into frames to obtain y. boom (n, x) and y ear (n, x), where yboom(n, x) is the time-domain signal of the x-th frame from the acquisition module on the microphone stick, y ear(n, x) represents the time-domain signal of the x-th frame from the acquisition module inside the earcup. The current sound energy data refers to the average energy value of the sound data in the current time frame. Here, n is the sampling point index within the frame (n=1, ..., N), x is the frame index, and N is the number of sampling points per frame (determined by the frame length and sampling rate; for example, with a 16kHz sampling rate and a 32ms frame length, N=512). Simultaneous processing of signals from both acquisition modules can be performed to avoid the need for detection and acquisition of a single set of data for speech acquisition, thereby improving the speech acquisition effect.
[0032] Step S12: When the current sound energy data does not match the preset user voice energy, determine the noise energy data based on the current sound energy data and the historical sound energy data of the previous frame. In this embodiment, after determining the current sound energy data of the current time frame, it is determined whether it is user speech based on the current sound energy data. For example, if the energy value of the current sound energy data is greater than a threshold, i.e., a preset user speech energy, it can be determined whether it is user speech; otherwise, it is determined to be environmental noise. For example, when it is determined that the current sound energy data does not match the preset user speech energy, i.e., it is determined to be environmental noise, the noise energy data is then determined based on the current sound energy data and the historical sound energy data of the previous frame. The noise energy data of the current frame is updated based on the time operation and the historical sound energy data of the previous frame to ensure the accuracy of the noise energy data of the current frame, thereby ensuring the accuracy of the final speech denoising. Here, the historical sound energy data of the previous frame refers to the noise energy data of the previous frame. For example, if the current time frame is the first frame, the historical sound energy data of the previous frame can be a predefined value. The noise energy data refers to the noise energy of the current frame, so as to facilitate subsequent denoising processing based on the noise energy.
[0033] For example, since the current sound energy data is actually determined by the first and second acquisition data, the current sound energy data is actually the sound energy data corresponding to the first and second acquisition data respectively. Then, based on these two sound energy data, a judgment process is executed to determine whether they match the preset user voice energy, and subsequent control is executed according to the matching result. This embodiment is only used as an example assuming that both sound energy data do not match the preset user voice energy. Of course, the energy value of the preset user voice energy can be appropriately varied depending on the distance between the acquisition module and the sound source.
[0034] Step S13: Determine the signal-to-noise ratio weight of the current time frame based on the noise energy data and the current sound energy data, and determine the target speech data based on the signal-to-noise ratio weight, the first acquisition data, and the second acquisition data.
[0035] In this embodiment, after determining the noise energy data, the signal-to-noise ratio (SNR) weight of the current time frame is determined based on the noise energy data and the current sound energy data. This weight is determined by the noise factor when synthesizing the final data from the two sets of acquired data. Then, the target speech data is determined based on the SNR weight, the first acquired data, and the second acquired data. This involves processing the proportions of the first and second acquired data in the overall signal based on the SNR weight to obtain the target speech data. The target speech data refers to the noise-reduced speech data. This allows for simultaneous speech acquisition using two acquisition modules to ensure effective speech acquisition and mutual noise reduction, thus guaranteeing the accuracy of the speech acquisition. In essence, the entire speech noise reduction process utilizes two acquisition modules on the headset for speech pickup and employs a weighted mixing method based on the SNR to handle the uplink speech noise reduction, avoiding the problems caused by switching between two microphones and improving the user experience during uplink voice calls.
[0036] In one embodiment, the current sound energy data includes first sound energy data of the current time frame in the first acquisition data and second sound energy data of the current time frame in the second acquisition data. The step of determining the current sound energy data of the current time frame based on the first acquisition data and the second acquisition data includes: Step S111: Determine the first current frame data of the current time frame in the first acquired data, and determine the first overall energy data of the current time frame based on the first current frame data. Use the quotient of the first overall energy data and the number of sampling points of the current time frame as the first sound energy data. Step S112: Determine the second current frame data of the current time frame in the second acquisition data, and determine the second overall energy data of the current time frame based on the second current frame data. Use the quotient of the second overall energy data and the number of sampling points of the current time frame as the second sound energy data.
[0037] In this embodiment, the current sound energy data includes the first sound energy data of the current time frame in the first acquisition data and the second sound energy data of the current time frame in the second acquisition data. That is, the entire speech denoising process is performed on the speech data of each frame. By determining the first current frame data of the current time frame in the first acquisition data, the first overall energy data of the current time frame is determined based on the first current frame data. Finally, the quotient obtained by dividing the first overall energy data of the current time frame by the number of sampling points of the current time frame is used as the first sound energy data. Among them, the first current frame data refers to the speech data acquired by the first acquisition module in the current time frame, the first overall energy data refers to the overall energy of the first current frame data, and the first sound energy data refers to the energy value of each sampling point in the current time frame. The calculation formula (1) for the first sound energy data is as follows: (1); Among them, P boom (x) represents the first sound energy data, N is the number of sampling points in the current time frame, and y boom (n, x) 2 The energy value of each sampling point. Simultaneously, the second current frame data of the current time frame in the second acquisition data is determined, and then the second overall energy data of the current time frame is determined based on the second current frame data. Finally, the quotient obtained by dividing the second overall energy data of the current time frame by the number of sampling points in the current time frame is used as the second sound energy data. Wherein, the second current frame data refers to the speech data acquired by the second acquisition module in the current time frame, the second overall energy data refers to the overall energy of the second current frame data, and the second sound energy data refers to the energy value of each sampling point in the current time frame. The calculation formula (2) for the second sound energy data is as follows: (2); Among them, P ear (x) represents the second sound energy data, N is the number of sampling points in the current time frame, and y ear (n, x) 2 This represents the energy value at each sampling point. This allows us to determine the current frame energy of each acquisition module, facilitating subsequent noise processing and ensuring the accuracy of speech acquisition from both acquisition modules.
[0038] For example, refer to Figure 4 , Figure 4This is a schematic diagram of the overall process of the speech noise reduction method of this application. It involves acquiring signals from a dual-pickup link, specifically the first acquisition data from the first acquisition module and the second acquisition data from the second acquisition module. The time-domain signals of the two acquisition data are then processed frame by frame to determine the current frame energy. Based on the current frame energy, it can be determined whether the energy is in the human voice frequency band. For example, a threshold can be used to determine if the current frame energy meets the requirements, or a VPU (Voice Processing Unit) can be used for VAD (Voice Activity Detection) to determine whether the signal in each frame contains user speech. The principle is to compare the current frame energy with a human voice energy threshold to determine if it is human speech. When it is determined to be human speech, the signal-to-noise ratio (SNR) is directly calculated, and the target speech data is obtained based on the mixed weights of the two sets of acquisition signals. Conversely, the noise energy is updated first, then the SNR is calculated, and finally, the mixed weights of the target speech data are determined. At this point, the two sets of collected data can be directly combined for voice acquisition processing to avoid the need to judge the usage status, thereby ensuring the voice acquisition effect. At the same time, noise reduction processing is performed on both sets of collected data to ensure the accuracy of the acquired voice.
[0039] Furthermore, based on the first and second embodiments of this application described above, a third embodiment of the speech noise reduction method of this application is proposed. In this embodiment, the noise energy data includes first noise energy data of the current time frame in the first acquisition data and second noise energy data of the current time frame in the second acquisition data. The step of determining the noise energy data based on the current sound energy data and the historical sound energy data of the previous frame includes: Step S121: Determine the first historical sound energy data of the first acquired data in the historical sound energy data of the previous frame, determine the first product between the first historical sound energy data and the preset first factor, determine the second product between the first sound energy data and the preset second factor in the current sound energy data, and take the sum of the first product and the second product as the first noise energy data, wherein the sum of the first factor and the second factor is one. Step S122: Determine the second historical sound energy data of the second acquired data in the historical sound energy data of the previous frame, determine the third product between the second historical sound energy data and the preset first factor, determine the fourth product between the second sound energy data in the current sound energy data and the preset second factor, and take the sum of the third product and the fourth product as the second noise energy data.
[0040] In this embodiment, when it is determined that the data is not human voice data, the noise energy is updated to ensure the accuracy of the final noise reduction. Taking the case where both sets of current sound energy data are not human voice data as an example, noise update processing is performed on both sets of current sound energy data separately. For example, the noise update process is as follows: First historical sound energy data is determined from the first acquired data in the historical sound energy data of the previous frame; a first product of the first historical sound energy data and a preset first factor is determined; a second product of the first sound energy data in the current sound energy data and a preset second factor is determined; and the sum of the first and second products is used as the first noise energy data. Here, the first historical sound energy data refers to the energy data of the first acquired data determined in the previous frame. Of course, the entire calculation process can also be to store the first historical sound energy data in the first buffer, store the preset first factor in the second buffer, and store the preset second factor in the third buffer. When it is determined that the step of determining the noise energy data based on the current sound energy data and the historical sound energy data of the previous frame needs to be executed, the data in the three buffers is directly called, and then processed with the first historical sound energy data through a multiplier and an adder to obtain the first noise energy data, that is, the energy data of the first acquisition module. Similarly, other calculation formulas in this application can also adopt the above buffer design processing method, which will not be described one by one here. Of course, the corresponding control algorithm can also be used to implement it, which will not be described one by one here. For example, the calculation formula (3) of the first noise energy data is as follows: (3); Where, N boom (x) represents the first noise energy data, α is the first factor with a value range of (0, 1), controlling the speed of noise estimation updates, 1-α is the second factor, and N boom (x-1) represents the first historical sound energy data. Similarly, the second noise energy data can be calculated based on the above method. By determining the historical sound energy data of the second acquisition data in the previous frame, and simultaneously determining the third product between the second historical sound energy data and the preset first factor, the fourth product between the second sound energy data and the preset second factor in the current sound energy data is determined, and the sum of the third and fourth products is taken as the second noise energy data. Here, the second historical sound energy data refers to the energy data of the second acquisition data determined in the previous frame. For example, the calculation formula (4) for the second noise energy data is as follows.
[0041] (4); Where, N ear (x) represents the second noise energy data, N ear(x-1) represents the second historical sound energy data. The first and second factors can be set to be the same or different according to the actual situation, and are not limited here. The entire noise update process can be adaptively updated based on the noise of the previous frame and the noise update factor to ensure the accuracy of subsequent noise processing, thereby ensuring the accuracy of speech acquisition.
[0042] In another embodiment, after the step of determining the current sound energy data of the current time frame based on the first acquisition data and the second acquisition data, the method includes: Step S113: When the current sound energy data matches the preset user voice energy, the historical sound energy data of the previous frame is used as noise energy data.
[0043] In this embodiment, when the VAD detects that the current frame is a speech frame, the noise energy can be kept unchanged, that is, the historical sound energy data of the previous frame can be used as the noise energy data. Of course, either one of the current sound energy data can be updated individually, and the other can be directly used as the historical sound energy data of the previous frame as the noise energy data. This can overcome the noise influence caused by the distance from the sound source and ensure the accuracy of the final noise processing.
[0044] Furthermore, based on the first, second, and / or third embodiments of this application described above, a fourth embodiment of the speech denoising method of this application is proposed. In this embodiment, the step of determining the signal-to-noise ratio weight of the current time frame based on noise energy data and current sound energy data includes: Step S131: Determine the first signal-to-noise ratio based on the first noise energy data in the noise energy data and the first sound energy data in the current sound energy data, and determine the first power ratio based on the first signal-to-noise ratio; Step S132: Determine the second signal-to-noise ratio (SNR) based on the second noise energy data in the noise energy data and the second sound energy data in the current sound energy data, determine the second power ratio based on the second SNR, and determine the SNR weight of the current time frame based on the second power ratio and the first power ratio.
[0045] In this embodiment, after determining the noise energy data of the two sets of signals, the signal-to-noise ratio (SNR) weights of each signal are determined based on the noise energy data. At this time, a first SNR is determined based on the first noise energy data in the noise energy data and the first sound energy data in the current sound energy data, and then a first power ratio is determined based on the first SNR. For example, the calculation formulas (5) for the first SNR and (6) for the first power ratio are as follows: (5); (6); Among them, SNR boom(x) represents the first signal-to-noise ratio, R boom (x) represents the first power ratio. Similarly, the second signal-to-noise ratio needs to be determined based on the second noise energy data in the noise energy data and the second sound energy data in the current sound energy data, and then the second power ratio is determined based on the second signal-to-noise ratio. For example, the calculation formula (7) for the second signal-to-noise ratio and the calculation formula (8) for the second power ratio are as follows: (5); (6); Among them, SNR ear (x) represents the second signal-to-noise ratio, R ear (x) represents the second power ratio. Based on the two power ratios, the final weight allocation for the two sets of signals can be determined, and the weight allocation can be adaptively performed based on noise to ensure the accuracy of the two sets of signal processing.
[0046] Furthermore, the signal-to-noise ratio (SNR) weights include a first SNR weight for the first acquired data and a second SNR weight for the second acquired data. The step of determining the SNR weight for the current time frame based on the second power ratio and the first power ratio includes: Step S1321: Determine the first signal-to-noise ratio weight based on the second power ratio, the first power ratio, and the preset time constant, and determine the second signal-to-noise ratio weight corresponding to the first signal-to-noise ratio weight, wherein the sum of the second signal-to-noise ratio weight and the first signal-to-noise ratio weight is one.
[0047] In this embodiment, after determining the second power ratio and the first power ratio, a first signal-to-noise ratio (SNR) weight is determined based on the second power ratio, the first power ratio, and a preset time constant. A corresponding second SNR weight is then determined based on the first SNR weight, where the sum of the second SNR weight and the first SNR weight is 1. For example, the calculation formula (9) for the first SNR weight and the calculation formula (10) for the second SNR weight are as follows: (9); (10); Among them, w ear As the second signal-to-noise ratio weight, w boom The first signal-to-noise ratio weight can be used to perform the final signal synthesis based on the two signal-to-noise ratio weights, ensuring the accuracy of speech acquisition.
[0048] Furthermore, based on the first, second, third, and / or fourth embodiments of this application described above, a fifth embodiment of the speech denoising method of this application is proposed. In this embodiment, the step of determining the target speech data based on the signal-to-noise ratio weight, the first acquired data, and the second acquired data includes: Step a: Determine the first target speech signal based on the first signal-to-noise ratio weight in the signal-to-noise ratio weight and the first acquired data, and determine the second target speech signal based on the second signal-to-noise ratio weight in the signal-to-noise ratio weight and the second acquired data; Step b: Determine the target speech data based on the first target speech signal and the second target speech signal.
[0049] In this embodiment, after determining the signal-to-noise ratio (SNR) weights of the two sets of signals, the final target speech data is determined based on the SNR weights. The calculation method for the target speech data is shown in (11): (11); Among them, S out (n, x) represents the target speech data, which is ultimately output through the DSP. Since this is the final processing of the speech, it is necessary to perform weighted comprehensive processing on the first and second target speech signals. The weighted comprehensive processing method can be a comprehensive selection of the two signals in the time domain and frequency domain, or it can be a simple addition processing on the time domain signals. This is not limited here.
[0050] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the speech noise reduction method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0051] This application also provides a speech noise reduction system, which includes a controller, a first acquisition module, and a second acquisition module. The controller is connected to the first and second acquisition modules as described above. Figure 5 The controller includes: The speech processing module is used to acquire first acquisition data from the first acquisition module and second acquisition data from the second acquisition module, and to determine target speech data based on the first acquisition data and the second acquisition data.
[0052] The controller provided in this application employs the speech denoising method described in the above embodiments, aiming to solve the technical problem of poor speech acquisition performance. Compared with the prior art, the beneficial effects of the controller provided in this application are the same as those of the speech denoising method provided in the above embodiments, and other technical features of the controller are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0053] This application provides a speech noise reduction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the speech noise reduction method in the first embodiment described above.
[0054] The following is for reference. Figure 6 The diagram illustrates a structural schematic suitable for implementing the speech noise reduction device in the embodiments of this application. The speech noise reduction device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The voice noise reduction device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0055] like Figure 6 As shown, the speech noise reduction device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the speech noise reduction device. The processing unit 1001, the ROM 1002, and the RAM 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the speech noise reduction device to communicate wirelessly or wiredly with other speech noise reduction devices to exchange data. Although speech noise reduction devices with various systems are shown in the figures, it should be understood that it is not required to implement or possess all of the systems shown. More or fewer systems may be implemented alternatively.
[0056] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0057] The speech noise reduction device provided in this application employs the speech noise reduction method in the above embodiments, aiming to solve the technical problem of poor speech acquisition effect. Compared with the prior art, the beneficial effects of the speech noise reduction device provided in this application are the same as those of the speech noise reduction method provided in the above embodiments, and other technical features in the speech noise reduction device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0058] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0059] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0060] This application provides a computer storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the speech noise reduction method in the above embodiments.
[0061] The computer storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0062] The aforementioned computer storage medium may be included in the speech noise reduction device; or it may exist independently and not be assembled into the speech noise reduction device.
[0063] The aforementioned computer storage medium carries one or more programs, which, when executed by the speech noise reduction device, cause the speech noise reduction device to perform the following: Acquire the first acquisition data from the first acquisition module and the second acquisition data from the second acquisition module, and determine the target speech data based on the first acquisition data and the second acquisition data.
[0064] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0065] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation that may be implemented in systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0066] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0067] The computer storage medium provided in this application stores computer-readable program instructions (i.e., computer programs) for executing the above-described speech denoising method, aiming to solve the technical problem of poor speech acquisition results. Compared with the prior art, the beneficial effects of the computer storage medium provided in this application are the same as the beneficial effects of the speech denoising method provided in the above embodiments, and will not be repeated here.
[0068] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the speech noise reduction method described above.
[0069] The computer program product provided in this application aims to solve the technical problem of poor voice acquisition results. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the voice noise reduction method provided in the above embodiments, and will not be repeated here.
[0070] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A speech noise reduction method, characterized in that, The speech denoising method is applied to a speech denoising system, which includes a first acquisition module and a second acquisition module. The speech denoising method includes: Acquire the first acquisition data from the first acquisition module and the second acquisition data from the second acquisition module, and determine the target speech data based on the first acquisition data and the second acquisition data.
2. The speech noise reduction method as described in claim 1, characterized in that, The step of determining the target speech data based on the first acquired data and the second acquired data includes: The current sound energy data for the current time frame is determined based on the first and second collected data. When the current sound energy data does not match the preset user voice energy, noise energy data is determined based on the current sound energy data and the historical sound energy data of the previous frame. The signal-to-noise ratio (SNR) weight of the current time frame is determined based on the noise energy data and the current sound energy data, and the target speech data is determined based on the SNR weight, the first acquired data, and the second acquired data.
3. The speech noise reduction method as described in claim 2, characterized in that, The current sound energy data includes the first sound energy data of the current time frame in the first acquisition data and the second sound energy data of the current time frame in the second acquisition data. The step of determining the current sound energy data of the current time frame based on the first acquisition data and the second acquisition data includes: Determine the first current frame data of the current time frame in the first collected data, and determine the first overall energy data of the current time frame based on the first current frame data. Use the quotient of the first overall energy data and the number of sampling points of the current time frame as the first sound energy data. The second current frame data of the current time frame in the second acquisition data is determined, and the second overall energy data of the current time frame is determined based on the second current frame data. The quotient of the second overall energy data and the number of sampling points of the current time frame is taken as the second sound energy data.
4. The speech noise reduction method as described in claim 2, characterized in that, The noise energy data includes the first noise energy data of the current time frame in the first acquired data and the second noise energy data of the current time frame in the second acquired data. The step of determining the noise energy data based on the current sound energy data and the historical sound energy data of the previous frame includes: The first historical sound energy data of the first collected data in the historical sound energy data of the previous frame is determined, the first product between the first historical sound energy data and the preset first factor is determined, the second product between the first sound energy data and the preset second factor in the current sound energy data is determined, and the sum of the first product and the second product is taken as the first noise energy data, wherein the sum of the first factor and the second factor is one; The second historical sound energy data of the second acquired data in the historical sound energy data of the previous frame is determined, the third product between the second historical sound energy data and the preset first factor is determined, the fourth product between the second sound energy data in the current sound energy data and the preset second factor is determined, and the sum of the third product and the fourth product is taken as the second noise energy data.
5. The speech noise reduction method as described in claim 2, characterized in that, The step of determining the signal-to-noise ratio weight of the current time frame based on the noise energy data and the current sound energy data includes: A first signal-to-noise ratio is determined based on the first noise energy data in the noise energy data and the first sound energy data in the current sound energy data, and a first power ratio is determined based on the first signal-to-noise ratio; A second signal-to-noise ratio (SNR) is determined based on the second noise energy data in the noise energy data and the second sound energy data in the current sound energy data, and a second power ratio is determined based on the second SNR. The SNR weight of the current time frame is determined based on the second power ratio and the first power ratio.
6. The speech noise reduction method as described in claim 5, characterized in that, The signal-to-noise ratio (SNR) weights include a first SNR weight for the first acquired data and a second SNR weight for the second acquired data. The step of determining the SNR weight of the current time frame based on the second power ratio and the first power ratio includes: The first signal-to-noise ratio weight is determined based on the second power ratio, the first power ratio, and a preset time constant, and the second signal-to-noise ratio weight corresponding to the first signal-to-noise ratio weight is determined, wherein the sum of the second signal-to-noise ratio weight and the first signal-to-noise ratio weight is one.
7. The speech noise reduction method as described in claim 2, characterized in that, The step of determining the target speech data based on the signal-to-noise ratio weight, the first acquired data, and the second acquired data includes: A first target speech signal is determined based on the first signal-to-noise ratio weight in the signal-to-noise ratio weight and the first acquired data, and a second target speech signal is determined based on the second signal-to-noise ratio weight in the signal-to-noise ratio weight and the second acquired data; Target speech data is determined based on the first target speech signal and the second target speech signal.
8. A speech noise reduction system, characterized in that, The speech noise reduction system includes a controller, a first acquisition module, and a second acquisition module. The controller is connected to the first acquisition module and the second acquisition module. The controller includes: The speech processing module is used to acquire the first acquisition data of the first acquisition module and the second acquisition data of the second acquisition module, and determine the target speech data based on the first acquisition data and the second acquisition data.
9. A voice noise reduction device, characterized in that, The speech noise reduction device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the speech noise reduction method as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that, The computer storage medium stores a speech denoising program, wherein when the speech denoising program is executed by a processor, it implements the steps of the speech denoising method as described in any one of claims 1 to 7.