Segmented energy-weighted noisy speech generation method, system, medium, and device

CN120895018BActive Publication Date: 2026-09-25广州广哈通信股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511108982.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2026-09-25
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

[0003]本申请提供了一种分段能量加权的带噪语音生成方法、系统、介质及设备,能够解决现有技术中生成的带噪语音不符合人耳对非稳态噪声的听觉感知特性的问题

Benefits of technology

;其中,表示带噪语音,表示原始人声音频。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895018B_ABST
    Figure CN120895018B_ABST
Patent Text Reader

Abstract

The application discloses a segmented energy weighted noisy speech generation method, system, medium and equipment, and belongs to the fields of voice communication and audio processing.The method is as follows: original human voice and noise audio of equal length are acquired, and the energy weights of each window of the two are calculated according to a preset signal window; the weighted energy equation is constructed by combining the energy weights of each window and the sampling point energy, and the first weighted energy of the human voice and the second weighted energy of the noise are obtained; the third weighted energy of the noise to be added is constructed and calculated according to the signal amplitude scaling coefficient of the original noise and the noise to be added. The amplitude scaling coefficient is solved through a signal-to-noise ratio equation by using the target signal-to-noise ratio, the first and third weighted energies, the original noise is adjusted to generate the noise to be added, and the noisy speech meeting the target signal-to-noise ratio is obtained by mixing the original human voice. Therefore, by implementing the application, the problem that the noisy speech generated in the prior art does not meet the auditory perception characteristics of the human ear to non-steady-state noise can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of voice communication and audio processing, and relates to a method, system, medium and device for generating noisy speech by segmented energy weighting. Background Technology

[0002] In the fields of voice communication and audio processing, the synthesis of noisy speech is a crucial step in evaluating the performance of noise reduction algorithms and training speech models. Its core requirement is to generate noisy speech data that is consistent with the listening experience of real-world environments. In existing technologies, traditional noisy speech generation methods determine the noise scaling ratio by calculating the average power ratio (signal-to-noise ratio) of the signal and noise. However, this method assumes that the noise is a stationary signal, which cannot adapt to non-stationary noise whose intensity changes over time in real-world scenarios. Furthermore, some improved schemes, such as the robust signal-to-noise ratio estimation technique based on the non-negative frequency-weighted energy operator of the derivative envelope of the speech signal proposed by Shome et al., employ a frequency-domain segmented weighting strategy, requiring adjustments to the weighting coefficients for different scenarios. This results in high computational complexity and poor versatility, ultimately leading to a significant deviation between the perceived noisy speech and the theoretical signal-to-noise ratio. Summary of the Invention

[0003] This application provides a segmented energy-weighted method, system, medium, and device for generating noisy speech, which can solve the problem that the noisy speech generated in the prior art does not conform to the auditory perception characteristics of the human ear for non-steady-state noise.

[0004] To achieve the above objectives, in a first aspect, the present invention provides a segmented energy-weighted noisy speech generation method, comprising: Obtain original human voice audio and original noise audio of equal duration; Based on a preset signal window, the energy weight of the audio signal in each signal window is calculated on the original human voice audio and the original noise audio respectively. Based on the energy weights and sampling point energy of the audio signals in each signal window, a weighted energy calculation equation is constructed, and the first weighted speech energy of the original human voice audio and the second weighted speech energy of the original noise audio are calculated respectively. Based on the preset signal amplitude scaling factor between the original noise audio and the noise to be added, and combined with the weighted energy calculation equation, a third weighted energy calculation equation for the noise to be added is constructed, and the third weighted speech energy of the noise to be added is calculated. The signal amplitude scaling factor is solved in the preset signal-to-noise ratio calculation equation by combining the preset target signal-to-noise ratio, the first weighted speech energy, and the third weighted speech energy. Based on the signal amplitude scaling factor, the original noise audio is adjusted to generate noise audio to be added, and combined with the original human voice audio to generate noisy speech that meets the target signal-to-noise ratio.

[0005] Compared with existing technologies, the embodiments of this application have the following beneficial effects: by acquiring original human voice and noise audio of equal duration, the synchronicity of subsequent energy calculation and noise addition processing is ensured, avoiding energy deviation caused by duration differences; based on a preset signal window, the energy weight of each window is calculated, and the energy proportion quantification is used to achieve differentiated weighting of speech activity segments (high energy) and silent segments (low energy), so that noise energy adjustment focuses on the speech activity periods that are sensitive to the human ear; a weighted energy equation is constructed to obtain the first and second weighted speech energy, and the ability to capture temporal energy distribution is refined through segmented energy characterization, adapting to the rapid energy change characteristics of non-steady-state noise (such as sudden and fluctuating noise); the third addition of noise is derived by combining the amplitude scaling coefficient. The weighted energy equation directly links noise adjustment with piecewise weighted energy, ensuring that the adjustment ratio of noise energy in each time window dynamically matches the speech energy. By solving for the scaling factor through the target signal-to-noise ratio and mixing the speech, noisy speech is finally generated, forming a closed loop of "perceptual needs - energy calculation - intensity control". This makes the noise energy distribution accurately match the time-varying perception characteristics of the human ear to non-steady-state noise. The synergistic effect of each step forms a complete technical chain of time-domain alignment → dynamic weighting → resolution improvement → quantization control → closed-loop control. Ultimately, the distribution law of noise energy in the generated noisy speech over time is consistent with the perception law of the human auditory system to non-steady-state noise, solving the technical contradiction of "achieving the target signal-to-noise ratio but hearing distortion" in traditional methods.

[0006] In some embodiments of the first aspect of this application, the step of calculating the energy weight of the audio signal for each signal window based on a preset signal window, respectively, on the original human voice audio and the original noise audio, includes: In the time domain, the original human voice audio and the original noise audio are divided into windows, and the energy weight of the corresponding audio signal is calculated for each signal window; the calculation formula is as follows: ; ;in, and These represent the energies of the original human voice signal and the original noise signal in the k-th window, respectively. and These represent the total energy of all windows for the original human voice audio and the original noise audio, respectively. and These represent the energy weights of the original human voice signal and the original noise signal in the k-th window, respectively.

[0007] Compared with the prior art, the above embodiments have the following beneficial effects: by dividing the audio into windows in the time domain and calculating the energy weights, and using the window energy ratio to quantify the weight values, dynamic focusing on the speech activity segment (high energy window) is achieved, the interference of the silent segment (low energy window) on the overall energy calculation is reduced, the time domain specificity of energy weight allocation is improved, and accurate weight parameter support is provided for the subsequent weighted energy equation.

[0008] In some embodiments of the first aspect of this application, the step of constructing a weighted energy calculation equation based on the energy weights and sampling point energy corresponding to the audio signals in each signal window, and calculating the first weighted speech energy of the original human voice audio and the second weighted speech energy of the original noise audio respectively, includes: Based on the energy weights and sampling point energy corresponding to each signal window, a first weighted energy calculation equation is constructed and solved to obtain the first weighted speech energy. The first weighted energy calculation equation is expressed as follows: ;in, This represents the first weighted speech energy of the original human voice audio. This represents the energy of the sampling point in window k where the absolute value of the original human voice signal is maximum. This represents the energy weight of the original human voice signal in the k-th window; Based on the energy weights and sampling point energy corresponding to each signal window, a second weighted energy calculation equation is constructed and solved to obtain the second weighted speech energy. The second weighted energy calculation equation is expressed as follows: ;in, This represents the second-weighted speech energy of the original noisy audio. This represents the energy of the sampling point in window k where the absolute value of the original noise signal is maximum. The value represents the energy weight of the original noise signal in the k-th window, and N represents the total number of signal windows for the original human voice audio and the total number of signal windows for the original noise audio.

[0009] Compared with the prior art, the above embodiments have the following beneficial effects: a weighted energy calculation equation is constructed based on energy weight and sampling point energy. By summing the product of the maximum energy of the window and the corresponding weight, the proportion of high-energy speech periods in the overall energy calculation is strengthened, the sensitivity of energy representation to the temporal changes of speech signals is improved, and the problem of loss of non-steady-state noise energy features caused by traditional full-time-domain average energy calculation is avoided.

[0010] In some embodiments of the first aspect of this application, the step of constructing a third weighted energy calculation equation for the noise to be added based on a preset signal amplitude scaling factor between the original noise audio and the noise to be added, combined with the weighted energy calculation equation, and calculating the third weighted speech energy of the noise to be added, includes: The numerical relationship between the energy of the sampling point with the maximum absolute value of the amplitude of the original noise signal within the signal window and the maximum signal amplitude is expressed as follows: ;in This represents the maximum signal amplitude of the original noise signal within the signal window k; The relationship between the maximum signal amplitude of the original noise audio and the noise to be added is expressed as follows: Where m represents the signal amplitude scaling factor between the original noise audio and the noise to be added. This represents the maximum signal amplitude of the noise to be added within the signal window k. The energy weighting between the original noise audio and the noise to be added is represented as follows: ;in, This represents the energy weight of the noise signal to be added in the k-th window; Combining the above three equations, the third weighted energy calculation equation for the noise to be added is constructed and solved to obtain the third weighted speech energy, as shown below: ; in, This represents the third weighted speech energy to which noise is to be added.

[0011] Compared with the prior art, the above embodiments have the following beneficial effects: by establishing the amplitude scaling relationship between the original noise and the noise to be added, the square term of the amplitude scaling factor is directly introduced into the noise energy calculation, realizing the physical dimension conversion from amplitude to energy without introducing additional parameters, ensuring the mathematical correlation between the noise energy adjustment process and the amplitude scaling factor, while maintaining the consistency of the weighting parameters in the noise energy calculation, simplifying the mapping logic from amplitude adjustment to energy control, and providing a rigorous mathematical basis for the subsequent solution of the scaling factor based on the target signal-to-noise ratio.

[0012] In some embodiments of the first aspect of this application, the step of solving for the signal amplitude scaling factor in a preset signal-to-noise ratio calculation equation by combining a preset target signal-to-noise ratio, the first weighted speech energy, and the third weighted speech energy includes: The signal-to-noise ratio calculation equation is expressed as follows: ;in, Indicates the signal-to-noise ratio. This represents the first weighted speech energy. This represents the third-weighted speech energy; In the signal-to-noise ratio calculation equation, given the target signal-to-noise ratio, and substituting it into the expression for the third weighted speech energy, the signal amplitude scaling factor is obtained, as follows: .

[0013] Compared with the prior art, the above embodiments have the following beneficial effects: by substituting the target signal-to-noise ratio into the weighted energy equation to solve for the scaling factor, and by transforming the nonlinear relationship between the signal-to-noise ratio and the energy ratio into a linear equation through logarithmic transformation, the scaling factor can be solved analytically, ensuring that the noise intensity control accuracy meets the target signal-to-noise ratio requirements.

[0014] In some embodiments of the first aspect of this application, the step of adjusting the original noisy audio according to the signal amplitude scaling factor to generate a noise-to-add audio, and combining it with the original human voice audio to generate noisy speech that meets the target signal-to-noise ratio, includes: The signal relationship between the original noise audio and the noise to be added is expressed as follows: ;in, This represents the noise to be added, and m represents the signal amplitude scaling factor. Represents the original noise audio; Substituting the expression for the signal amplitude scaling factor, the generated audio with added noise is shown below: ; Based on the audio to be noised, and combined with the original human voice audio, noisy speech that meets the target signal-to-noise ratio is generated, as follows: ;in, This indicates noisy speech. This represents the original human voice audio.

[0015] Compared with the prior art, the above embodiments have the following beneficial effects: the amplitude of the noise to be added is adjusted according to the scaling factor and mixed with the clean speech, and the noise adjustment and speech synthesis process is directly quantified by mathematical formulas, ensuring the consistency between the theoretically calculated value of the signal-to-noise ratio of the noisy speech and the actual auditory perception.

[0016] In some embodiments of the first aspect of this application, the size of the signal window is 10 to 30 ms.

[0017] Compared with the prior art, the above embodiments have the following beneficial effects: considering the diversity of different audio signals and the complexity of noise energy distribution, the signal window size is limited to 10-30ms. This range can balance time domain resolution and computational stability: if the window is too small, it will cause drastic energy fluctuations, and if the window is too large, it will mask the short-term energy changes of non-steady-state noise; it can effectively adapt to the human ear's perception time scale of non-steady-state noise.

[0018] Secondly, the present invention also provides a segmented energy-weighted noisy speech generation system, comprising: a data acquisition module, a weight calculation module, a weighted energy calculation module, a third weighted energy calculation module, a solution module, and a result output module; The data acquisition module is used to acquire original human voice audio and original noise audio of equal duration. The weight calculation module is used to calculate the energy weight of the audio signal of each signal window based on the preset signal window, respectively, on the original human voice audio and the original noise audio. The weighted energy calculation module is used to construct a weighted energy calculation equation based on the energy weight and sampling point energy of the audio signal in each signal window, and to calculate the first weighted speech energy of the original human voice audio and the second weighted speech energy of the original noise audio respectively. The third weighted energy calculation module is used to construct a third weighted energy calculation equation for the noise to be added based on a preset signal amplitude scaling factor between the original noise audio and the noise to be added, combined with the weighted energy calculation equation, and to calculate the third weighted speech energy of the noise to be added. The solution module is used to solve the signal amplitude scaling factor in a preset signal-to-noise ratio calculation equation by combining the preset target signal-to-noise ratio, the first weighted speech energy and the third weighted speech energy. The result output module is used to adjust the original noise audio according to the signal amplitude scaling factor, generate noise audio to be added, and combine it with the original human voice audio to generate noisy speech that meets the target signal-to-noise ratio.

[0019] Compared with existing technologies, the above embodiments of this application have the following beneficial effects: by acquiring original human voice and noise audio of equal duration, the synchronicity of subsequent energy calculation and noise addition processing is ensured, avoiding energy deviation caused by duration differences; based on a preset signal window, the energy weight of each window is calculated, and the energy proportion quantification is used to achieve differentiated weighting of speech activity segments (high energy) and silent segments (low energy), so that noise energy adjustment focuses on the speech activity periods that are sensitive to the human ear; a weighted energy equation is constructed to obtain the first and second weighted speech energy, and the ability to capture temporal energy distribution is refined through segmented energy characterization, adapting to the rapid energy change characteristics of non-steady-state noise (such as sudden and fluctuating noise); the third weighted energy of the noise to be added is derived by combining the amplitude scaling coefficient. The weighted energy equation directly links noise adjustment with piecewise weighted energy, ensuring that the adjustment ratio of noise energy in each time window dynamically matches the speech energy. By solving for the scaling factor through the target signal-to-noise ratio and mixing the speech, noisy speech is finally generated, forming a closed loop of "perceptual needs - energy calculation - intensity control". This makes the noise energy distribution accurately match the time-varying perception characteristics of the human ear to non-steady-state noise. The synergistic effect of each step forms a complete technical chain of time-domain alignment → dynamic weighting → resolution improvement → quantization control → closed-loop control. Ultimately, the distribution law of noise energy in the generated noisy speech over time is consistent with the perception law of the human auditory system to non-steady-state noise, solving the technical contradiction of "achieving the target signal-to-noise ratio but hearing distortion" in traditional methods.

[0020] Thirdly, the present invention also provides a segmented energy-weighted noisy speech generation device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it implements the steps of any segmented energy-weighted noisy speech generation method of the present invention.

[0021] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the segmented energy-weighted noisy speech generation methods of the present invention. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating a segmented energy-weighted noisy speech generation method provided in some embodiments of the present invention.

[0023] Figure 2 This is a schematic diagram of a segmented energy-weighted noisy speech generation system provided in some embodiments of the present invention.

[0024] Figure 3 : This is a structural diagram of a segmented energy-weighted noisy speech generation device provided in some embodiments of the present invention.

[0025] Figure 4 This is a schematic diagram illustrating the relationship between window size and voltage level in some embodiments of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] Example 1: Please refer to Figure 1 To address the problem that noisy speech generated in existing technologies does not conform to the auditory perception characteristics of the human ear for non-steady-state noise, an embodiment of the present invention provides a segmented energy-weighted noisy speech generation method, comprising steps S1 to S6: Step S1: Obtain original human voice audio and original noise audio of equal duration.

[0028] In practice, to ensure the quality of the final noisy speech, the original human voice audio and the original noise audio should be as clean as possible.

[0029] Step S2: Based on the preset signal window, calculate the energy weight of the audio signal for each signal window on the original human voice audio and the original noise audio respectively.

[0030] Furthermore, step S2 can be implemented through the following preferred embodiment, including step S21, as follows: S21: In the time domain, the original human voice audio and the original noise audio are divided into windows respectively, and the energy weight of the corresponding audio signal is calculated in each signal window; wherein, the calculation formula is as follows: ; ;in, and These represent the energies of the original human voice signal and the original noise signal in the k-th window, respectively. and These represent the total energy of all windows for the original human voice audio and the original noise audio, respectively. and These represent the energy weights of the original human voice signal and the original noise signal in the k-th window, respectively.

[0031] In this preferred embodiment, by dividing the audio into windows in the time domain and calculating energy weights, the weight values ​​are quantified using the window energy ratio, thereby achieving dynamic focusing on the speech activity segment (high energy window), reducing the interference of the silent segment (low energy window) on the overall energy calculation, improving the time domain specificity of energy weight allocation, and providing accurate weight parameter support for the subsequent weighted energy equation.

[0032] Furthermore, the size of the signal window is 10–30 ms.

[0033] In this preferred embodiment, considering the diversity of different audio signals and the complexity of noise energy distribution, the signal window size is limited to 10-30ms. This range can balance time domain resolution and computational stability: too small a window will cause drastic energy fluctuations, while too large a window will mask the short-term energy changes of non-steady-state noise; it can effectively adapt to the human ear's perception time scale of non-steady-state noise.

[0034] Step S3: Based on the energy weights and sampling point energy of the audio signals in each signal window, construct a weighted energy calculation equation, and calculate the first weighted speech energy of the original human voice audio and the second weighted speech energy of the original noise audio respectively.

[0035] Furthermore, step S3 can be implemented through the following preferred embodiments, including steps S31-S32, as follows: S31: Based on the energy weights and sampling point energy corresponding to each signal window, a first weighted energy calculation equation is constructed and solved to obtain the first weighted speech energy. The first weighted energy calculation equation is expressed as follows: ;in, This represents the first weighted speech energy of the original human voice audio. This represents the energy of the original human voice signal in window k at the sampling point where the absolute value of the amplitude is the largest (i.e., the square of the amplitude at that point). This represents the energy weight of the original human voice signal in the k-th window; S32: Based on the energy weights and sampling point energy corresponding to each signal window, a second weighted energy calculation equation is constructed and solved to obtain the second weighted speech energy on the original noisy audio; wherein, the second weighted energy calculation equation is expressed as follows: ;in, This represents the second-weighted speech energy of the original noisy audio. This represents the energy of the sampling point in window k where the absolute value of the original noise signal is maximum. The value represents the energy weight of the original noise signal in the k-th window, and N represents the total number of signal windows for the original human voice audio and the total number of signal windows for the original noise audio.

[0036] In this preferred embodiment, a weighted energy calculation equation is constructed based on energy weights and sampling point energy. By summing the product of the maximum energy of the window and the corresponding weight, the proportion of high-energy speech periods in the overall energy calculation is strengthened, the sensitivity of energy representation to temporal changes in speech signals is improved, and the problem of loss of non-steady-state noise energy features caused by traditional full-time-domain average energy calculation is avoided.

[0037] In practice, the original audio is first divided into N short-time analysis windows in the time domain. The corresponding weight is calculated in each window. Then, in order to more accurately describe the salient region of the speech signal, the concept of weighted speech energy is introduced to represent the intensity characteristics of the entire speech signal. The calculation method is as described in step S3 above.

[0038] In order to study the impact of window size on the subsequent calculation results of this application, this application designs a level calculation formula. If the amplitude of the audio sampling points is normalized to [-1.0, 1.0], according to the mathematical characteristics of the energy weight in step S2 and the weighted energy formula proposed in step S3: , ; From the two formulas above, The maximum value is 1.

[0039] The formula for calculating the level is: ;in, Indicates the amplitude of the speech signal. This represents the maximum amplitude of the audio sample (the maximum energy amplitude of the sample value may be different for different digital sampling precisions; for example, the maximum amplitude of 16-bit sampling is (2^16) / 2. Here, due to normalization, the maximum amplitude is 1).

[0040] Will The maximum value of 1 is used as , Replace with The formula for calculating the voltage level can be obtained as follows: ;in The value is affected by the window size (because) and (Both are affected by the window size).

[0041] During the experiment, different types of audio files (various steady-state and non-steady-state noises, including white noise, machine noise, stool pulling sounds, coughing sounds, keyboard sounds, human voices, etc.) were selected, and graphing software was used to analyze the window size and level. Relationships, such as Figure 4The diagram illustrates the relationship between window size and audio level. For various sounds, when the window size is greater than approximately 80ms, the audio level changes significantly with the window size. However, if the window size is too small, less than approximately 2ms, a jagged pattern of level variation appears. Overall, the calculated audio level for various sounds is relatively stable within a window size range of 10ms to 50ms.

[0042] In addition, there is a concept called auditory masking, which refers to the temporary inability of the human ear to hear the subsequent weaker sound after a loud sound. This phenomenon includes two situations: (1) Forward masking: The stronger sound will mask the weaker sound within about 50ms to 200ms after it ends.

[0043] (2) Backward masking: A stronger sound will mask a weaker sound about 10 to 20 ms before it appears.

[0044] Based on the above experimental analysis and the auditory masking phenomenon, and considering the computational load, a window size of 10ms to 30ms can be selected as an optimal value, such as 20ms.

[0045] Step S4: Based on the preset signal amplitude scaling factor between the original noise audio and the noise to be added, and combined with the weighted energy calculation equation, construct the third weighted energy calculation equation for the noise to be added, and calculate the third weighted speech energy of the noise to be added.

[0046] Furthermore, step S4 can be implemented through the following preferred embodiment, including step S41, as follows: S41: The numerical relationship between the energy of the sampling point where the absolute value of the amplitude of the original noise signal is the maximum in the signal window and the maximum signal amplitude is expressed as follows: ;in This represents the maximum signal amplitude of the original noise signal within the signal window k; The relationship between the maximum signal amplitude of the original noise audio and the noise to be added is expressed as follows: Where m represents the signal amplitude scaling factor between the original noise audio and the noise to be added. This represents the maximum signal amplitude of the noise to be added within the signal window k. The energy weighting between the original noise audio and the noise to be added is represented as follows: ;in, This represents the energy weight of the noise signal to be added in the k-th window; Combining the above three equations, the third weighted energy calculation equation for the noise to be added is constructed and solved to obtain the third weighted speech energy, as shown below: ; in, This represents the third weighted speech energy to which noise is to be added.

[0047] In this preferred embodiment, by establishing the amplitude scaling relationship between the original noise and the noise to be added, the square term of the amplitude scaling factor is directly introduced into the noise energy calculation. This achieves the physical dimension conversion from amplitude to energy without introducing additional parameters, ensuring the mathematical correlation between the noise energy adjustment process and the amplitude scaling factor. At the same time, it maintains the consistency of the weighting parameters in the noise energy calculation, simplifies the mapping logic from amplitude adjustment to energy control, and provides a rigorous mathematical basis for the subsequent solution of the scaling factor based on the target signal-to-noise ratio.

[0048] Step S5: Combine the preset target signal-to-noise ratio, the first weighted speech energy, and the third weighted speech energy to solve for the signal amplitude scaling factor in the preset signal-to-noise ratio calculation equation.

[0049] Furthermore, step S5 can be implemented through the following preferred embodiment, including step S51, as follows: S51: The signal-to-noise ratio calculation equation is expressed as follows: ;in, Indicates the signal-to-noise ratio. This represents the first weighted speech energy. This represents the third-weighted speech energy; In the signal-to-noise ratio calculation equation, given the target signal-to-noise ratio, and substituting it into the expression for the third weighted speech energy, the signal amplitude scaling factor is obtained, as follows: .

[0050] In this preferred embodiment, the target signal-to-noise ratio is substituted into the weighted energy equation to solve for the scaling factor. The nonlinear relationship between the signal-to-noise ratio and the energy ratio is transformed into a linear equation through logarithmic transformation, thereby realizing the analytical solution of the scaling factor and ensuring that the noise intensity control accuracy meets the target signal-to-noise ratio requirements.

[0051] Furthermore, compared to other signal-to-noise ratio (SNR) calculation methods, this application's weighted speech energy substitution method is adaptable to both steady-state and non-steady-state noise, is more sensitive to the sound-generating region, and has stronger suppression of silent regions. Specifically, for steady-state noise, since its energy is relatively stable in the time dimension, the SNR after segmented weighting is basically consistent with the traditional average SNR result. This application does not introduce additional bias under steady-state noise conditions and can accurately reflect the noise level. For non-steady-state noise, since this method pays more attention to the speech activity segment during the weighting process, the SNR calculation result can highlight the degree of noise influence during speech, enhance the adaptability to non-stationary background noise, and demonstrate good versatility and practicality. Moreover, the calculation process is simple and efficient, making it easy to apply in practice. Furthermore, based on this SNR calculation method, the synthesis formula for the audio to be noised is derived, which can generate noisy speech data that better conforms to the characteristics of human auditory perception, providing a more realistic and effective simulation environment for training and evaluating speech denoising models. The specific method is as follows: Step S6: Adjust the original noise audio according to the signal amplitude scaling factor to generate noise audio to be added, and combine it with the original human voice audio to generate noisy speech that meets the target signal-to-noise ratio.

[0052] Furthermore, step S6 can be implemented through the following preferred embodiments, including steps S61-S62, as follows: S61: The signal relationship between the original noise audio and the noise to be added is expressed as follows: ;in, This represents the noise to be added, and m represents the signal amplitude scaling factor. Represents the original noise audio; Substituting the expression for the signal amplitude scaling factor, the generated audio with added noise is shown below: ; S62: Based on the audio to be noised and combined with the original human voice audio, generate noisy speech that meets the target signal-to-noise ratio, as follows: ;in, This indicates noisy speech. This represents the original human voice audio.

[0053] In this preferred embodiment, the amplitude of the noise to be added is adjusted according to the scaling factor and mixed with the clean speech. The noise adjustment and speech synthesis process is directly quantified by mathematical formulas to ensure the consistency between the theoretically calculated value of the signal-to-noise ratio of the noisy speech and the actual auditory perception.

[0054] In summary, compared with the prior art, the above embodiments of this application have the following beneficial effects: by acquiring original human voice and noise audio of equal duration, the synchronicity of subsequent energy calculation and noise addition processing is ensured, avoiding energy deviation caused by duration differences; based on a preset signal window, the energy weight of each window is calculated, and the energy proportion quantification is used to achieve differentiated weighting of speech activity segments (high energy) and silent segments (low energy), so that noise energy adjustment focuses on the speech activity periods that are sensitive to the human ear; a weighted energy equation is constructed to obtain the first and second weighted speech energy, and the ability to capture temporal energy distribution is refined through segmented energy characterization, adapting to the rapid energy change characteristics of non-steady-state noise (such as sudden and fluctuating noise); the first and second weighted speech energy are derived by combining the amplitude scaling coefficient. The three-weighted energy equation directly links noise adjustment with piecewise weighted energy, ensuring that the adjustment ratio of noise energy in each time window dynamically matches the speech energy. By solving for the scaling factor through the target signal-to-noise ratio and mixing the speech, noisy speech is finally generated, forming a closed loop of "perceptual needs - energy calculation - intensity control". This makes the noise energy distribution accurately match the time-varying perception characteristics of the human ear to non-steady-state noise. The synergistic effect of each step forms a complete technical chain of time-domain alignment → dynamic weighting → resolution improvement → quantization control → closed-loop control. Ultimately, the distribution law of noise energy in the generated noisy speech over time is consistent with the perception law of the human auditory system to non-steady-state noise, solving the technical contradiction of "achieving the target signal-to-noise ratio but hearing distortion" in traditional methods.

[0055] Example 2: Please refer to Figure 2 Based on the same inventive concept, the present invention discloses a segmented energy weighted noisy speech generation system, comprising: a data acquisition module M1, a weight calculation module M2, a weighted energy calculation module M3, a third weighted energy calculation module M4, a solution module M5, and a result output module M6. The data acquisition module M1 is used to acquire original human voice audio and original noise audio of equal duration.

[0056] The weight calculation module M2 is used to calculate the energy weight of the audio signal of each signal window based on the preset signal window, respectively, on the original human voice audio and the original noise audio.

[0057] Furthermore, the weight calculation module M2 includes: a partitioning calculation unit; The partitioning calculation unit is used to partition the original human voice audio and the original noise audio into windows in the time domain, and calculate the energy weight of the corresponding audio signal in each signal window; the calculation formula is as follows: ; ;in, and These represent the energies of the original human voice signal and the original noise signal in the k-th window, respectively. and These represent the total energy of all windows for the original human voice audio and the original noise audio, respectively. and These represent the energy weights of the original human voice signal and the original noise signal in the k-th window, respectively.

[0058] In this preferred embodiment, by dividing the audio into windows in the time domain and calculating energy weights, the weight values ​​are quantified using the window energy ratio, thereby achieving dynamic focusing on the speech activity segment (high energy window), reducing the interference of the silent segment (low energy window) on the overall energy calculation, improving the time domain specificity of energy weight allocation, and providing accurate weight parameter support for the subsequent weighted energy equation.

[0059] Furthermore, the size of the signal window is 10–30 ms.

[0060] In this preferred embodiment, considering the diversity of different audio signals and the complexity of noise energy distribution, the signal window size is limited to 10-30ms. This range can balance time domain resolution and computational stability: too small a window will cause drastic energy fluctuations, while too large a window will mask the short-term energy changes of non-steady-state noise; it can effectively adapt to the human ear's perception time scale of non-steady-state noise.

[0061] The weighted energy calculation module M3 is used to construct a weighted energy calculation equation based on the energy weights and sampling point energy of the audio signals in each signal window, and to calculate the first weighted speech energy of the original human voice audio and the second weighted speech energy of the original noise audio respectively.

[0062] Furthermore, the weighted energy calculation module M3 includes: a first weighted energy calculation unit and a second weighted energy calculation unit; The first weighted energy calculation unit is used to construct a first weighted energy calculation equation and solve it to obtain the first weighted speech energy based on the energy weights and sampling point energy corresponding to each signal window on the original human voice audio. The first weighted energy calculation equation is expressed as follows: ;in, This represents the first weighted speech energy of the original human voice audio. This represents the energy of the sampling point in window k where the absolute value of the original human voice signal is maximum. This represents the energy weight of the original human voice signal in the k-th window; The second weighted energy calculation unit is used to construct a second weighted energy calculation equation and solve it to obtain the second weighted speech energy based on the energy weights corresponding to each signal window and the energy of the sampling points on the original noisy audio; wherein, the second weighted energy calculation equation is expressed as follows: ;in, This represents the second-weighted speech energy of the original noisy audio. This represents the energy of the sampling point in window k where the absolute value of the original noise signal is maximum. The value represents the energy weight of the original noise signal in the k-th window, and N represents the total number of signal windows for the original human voice audio and the total number of signal windows for the original noise audio.

[0063] In this preferred embodiment, a weighted energy calculation equation is constructed based on energy weights and sampling point energy. By summing the product of the maximum energy of the window and the corresponding weight, the proportion of high-energy speech periods in the overall energy calculation is strengthened, the sensitivity of energy representation to temporal changes in speech signals is improved, and the problem of loss of non-steady-state noise energy features caused by traditional full-time-domain average energy calculation is avoided.

[0064] The third weighted energy calculation module M4 is used to construct the third weighted energy calculation equation of the noise to be added based on the preset signal amplitude scaling coefficient between the original noise audio and the noise to be added, combined with the weighted energy calculation equation, and to calculate the third weighted speech energy of the noise to be added.

[0065] Furthermore, the third weighted energy calculation module M4 includes: a joint solution unit; The numerical relationship between the energy of the sampling point with the maximum absolute value of the amplitude of the original noise signal within the signal window and the maximum signal amplitude is expressed as follows: ;in This represents the maximum signal amplitude of the original noise signal within the signal window k; The relationship between the maximum signal amplitude of the original noise audio and the noise to be added is expressed as follows: Where m represents the signal amplitude scaling factor between the original noise audio and the noise to be added. This represents the maximum signal amplitude of the noise to be added within the signal window k. The energy weighting between the original noise audio and the noise to be added is represented as follows: ;in, This represents the energy weight of the noise signal to be added in the k-th window; The joint solving unit is used to combine the above three equations to construct the third weighted energy calculation equation for the noise to be added and solve for the third weighted speech energy, as shown below: ; in, This represents the third weighted speech energy of the noise to be added.

[0066] In this preferred embodiment, by establishing the amplitude scaling relationship between the original noise and the noise to be added, the square term of the amplitude scaling factor is directly introduced into the noise energy calculation. This achieves the physical dimension conversion from amplitude to energy without introducing additional parameters, ensuring the mathematical correlation between the noise energy adjustment process and the amplitude scaling factor. At the same time, it maintains the consistency of the weighting parameters in the noise energy calculation, simplifies the mapping logic from amplitude adjustment to energy control, and provides a rigorous mathematical basis for the subsequent solution of the scaling factor based on the target signal-to-noise ratio.

[0067] The solution module M5 is used to solve the signal amplitude scaling factor in a preset signal-to-noise ratio calculation equation by combining the preset target signal-to-noise ratio, the first weighted speech energy, and the third weighted speech energy.

[0068] Furthermore, the solution module M5 includes: a coefficient solution unit; The signal-to-noise ratio calculation equation is expressed as follows: ;in, Indicates the signal-to-noise ratio. This represents the first weighted speech energy. This represents the third-weighted speech energy; The coefficient solving unit is used to solve for the signal amplitude scaling coefficient in the signal-to-noise ratio calculation equation, given the target signal-to-noise ratio and substituting it into the expression for the third weighted speech energy, as follows: .

[0069] In this preferred embodiment, the target signal-to-noise ratio is substituted into the weighted energy equation to solve for the scaling factor. The nonlinear relationship between the signal-to-noise ratio and the energy ratio is transformed into a linear equation through logarithmic transformation, thereby realizing the analytical solution of the scaling factor and ensuring that the noise intensity control accuracy meets the target signal-to-noise ratio requirements.

[0070] The result output module M6 is used to adjust the original noise audio according to the signal amplitude scaling factor, generate noise audio to be added, and combine it with the original human voice audio to generate noisy speech that meets the target signal-to-noise ratio.

[0071] Furthermore, the result output module M6 includes: a substitution solution unit and a combination generation unit; The signal relationship between the original noise audio and the noise to be added is expressed as follows: ;in, This represents the noise to be added, and m represents the signal amplitude scaling factor. Represents the original noise audio; The substitution and solving unit is used to substitute the expression for the signal amplitude scaling factor to generate the audio to be noised, as shown below: ; The combined generation unit is used to generate noisy speech that meets the target signal-to-noise ratio based on the audio to be noised and the original human voice audio, as shown below: ;in, This indicates noisy speech. This represents the original human voice audio.

[0072] In this preferred embodiment, the amplitude of the noise to be added is adjusted according to the scaling factor and mixed with the clean speech. The noise adjustment and speech synthesis process is directly quantified by mathematical formulas to ensure the consistency between the theoretically calculated value of the signal-to-noise ratio of the noisy speech and the actual auditory perception.

[0073] In summary, compared with the prior art, the embodiments of this application have the following beneficial effects: by acquiring original human voice and noise audio of equal duration, the synchronicity of subsequent energy calculation and noise addition processing is ensured, avoiding energy deviation caused by duration differences; based on a preset signal window, the energy weight of each window is calculated, and the energy proportion quantification is used to achieve differentiated weighting of speech activity segments (high energy) and silent segments (low energy), so that noise energy adjustment focuses on the speech activity periods that are sensitive to the human ear; a weighted energy equation is constructed to obtain the first and second weighted speech energy, and the ability to capture temporal energy distribution is refined through segmented energy characterization, adapting to the rapid energy change characteristics of non-steady-state noise (such as sudden and fluctuating noise); the third weighted energy of the noise to be added is derived by combining the amplitude scaling coefficient. The weighted energy equation directly links noise adjustment with piecewise weighted energy, ensuring that the adjustment ratio of noise energy in each time window dynamically matches the speech energy. By solving for the scaling factor through the target signal-to-noise ratio and mixing the speech, noisy speech is finally generated, forming a closed loop of "perceptual needs - energy calculation - intensity control". This makes the noise energy distribution accurately match the time-varying perception characteristics of the human ear to non-steady-state noise. The synergistic effect of each step forms a complete technical chain of time-domain alignment → dynamic weighting → resolution improvement → quantization control → closed-loop control. Ultimately, the distribution law of noise energy in the generated noisy speech over time is consistent with the perception law of the human auditory system to non-steady-state noise, solving the technical contradiction of "achieving the target signal-to-noise ratio but hearing distortion" in traditional methods.

[0074] Example 3: Figure 3 A structural diagram of a segmented energy-weighted noisy speech generation device according to this application is presented. Figure 3 As shown, the segmented energy-weighted noisy speech generation device may include: a processor N1, a memory N2, a data interface N3, and a communication bus N4.

[0075] Wherein: processor N1, memory N2, and data interface N3 communicate with each other through communication bus N4; data interface N3 is used for data communication with other devices such as input devices or output devices; processor N1 is used to execute program N5, specifically it can execute the relevant steps in any of the above-mentioned segmented energy weighted noisy speech generation method embodiments.

[0076] Specifically, program N5 may include program code, which includes computer-executable instructions.

[0077] The processor N1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The segmented energy-weighted noisy speech generation device includes one or more processors, which may be processors of the same type, such as one or more CPUs, or processors of different types, such as one or more CPUs and one or more ASICs.

[0078] Memory N2 is used to store program N5. Memory N2 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage.

[0079] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Furthermore, the embodiments in this application are not directed to any particular programming language.

[0080] Example 4: This invention also provides a computer-readable storage medium storing at least one executable instruction that, when executed on a segmented energy-weighted noisy speech generation device / system, causes the segmented energy-weighted noisy speech generation device / system to perform one of the segmented energy-weighted noisy speech generation methods described in any of the above method embodiments.

[0081] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. Similarly, for the purpose of simplification and aiding understanding of one or more aspects of the invention, in the above description of exemplary embodiments of this application, various features of the embodiments are sometimes grouped together in a single embodiment, figure, or description thereof. The claims, which follow the detailed description, are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of this application.

[0082] Those skilled in the art will understand that the modules in the device of the embodiment can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiment can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components, except that at least some of such features and / or processes or units are mutually exclusive.

Claims

1. A segmented energy-weighted method for generating noisy speech, characterized in that, include: Obtain original human voice audio and original noise audio of equal duration; Based on a preset signal window, the energy weight of the audio signal in each signal window is calculated on the original human voice audio and the original noise audio respectively. Based on the energy weights and sampling point energy of the audio signals in each signal window, a weighted energy calculation equation is constructed, and the first weighted speech energy of the original human voice audio and the second weighted speech energy of the original noise audio are calculated respectively. Based on the preset signal amplitude scaling factor between the original noise audio and the noise to be added, and combined with the second weighted speech energy, a third weighted energy calculation equation for the noise to be added is constructed, and the third weighted speech energy of the noise to be added is calculated. The signal amplitude scaling factor is solved in the preset signal-to-noise ratio calculation equation by combining the preset target signal-to-noise ratio, the first weighted speech energy, and the third weighted speech energy. Based on the signal amplitude scaling factor, the original noise audio is adjusted to generate noise audio to be added, and combined with the original human voice audio to generate noisy speech that meets the target signal-to-noise ratio.

2. The segmented energy-weighted noisy speech generation method as described in claim 1, characterized in that, The method, based on a preset signal window, calculates the energy weight of the audio signal for each signal window on both the original human voice audio and the original noise audio, including: In the time domain, the original human voice audio and the original noise audio are divided into windows, and the energy weight of the corresponding audio signal is calculated for each signal window; the calculation formula is as follows: ; ;in, and These represent the energies of the original human voice signal and the original noise signal in the k-th window, respectively. and These represent the total energy of all windows for the original human voice audio and the original noise audio, respectively. and These represent the energy weights of the original human voice signal and the original noise signal in the k-th window, respectively.

3. The segmented energy-weighted noisy speech generation method as described in claim 2, characterized in that, The step involves constructing a weighted energy calculation equation based on the energy weights and sampling point energy of the audio signal in each signal window, and calculating the first weighted speech energy of the original human voice audio and the second weighted speech energy of the original noise audio, including: Based on the energy weights and sampling point energy corresponding to each signal window, a first weighted energy calculation equation is constructed and solved to obtain the first weighted speech energy. The first weighted energy calculation equation is expressed as follows: ;in, This represents the first weighted speech energy of the original human voice audio. This represents the energy of the sampling point in window k where the absolute value of the original human voice signal is maximum. This represents the energy weight of the original human voice signal in the k-th window; Based on the energy weights and sampling point energy corresponding to each signal window, a second weighted energy calculation equation is constructed and solved to obtain the second weighted speech energy. The second weighted energy calculation equation is expressed as follows: ;in, This represents the second-weighted speech energy of the original noisy audio. This represents the energy of the sampling point in window k where the absolute value of the original noise signal is maximum. The value represents the energy weight of the original noise signal in the k-th window, and N represents the total number of signal windows for the original human voice audio and the total number of signal windows for the original noise audio.

4. The segmented energy-weighted noisy speech generation method as described in claim 3, characterized in that, The process involves constructing a third weighted energy calculation equation for the noise to be added based on a preset signal amplitude scaling factor between the original noise audio and the noise to be added, combined with the second weighted speech energy, and calculating the third weighted speech energy of the noise to be added, including: The numerical relationship between the energy of the sampling point with the maximum absolute value of the amplitude of the original noise signal within the signal window and the maximum signal amplitude is expressed as follows: ;in This represents the maximum signal amplitude of the original noise signal within the signal window k; The relationship between the maximum signal amplitude of the original noise audio and the noise to be added is expressed as follows: Where m represents the signal amplitude scaling factor between the original noise audio and the noise to be added. This represents the maximum signal amplitude of the noise to be added within the signal window k. The energy weighting between the original noise audio and the noise to be added is represented as follows: ;in, This represents the energy weight of the noise signal to be added in the k-th window; Combining the above three equations, the third weighted energy calculation equation for the noise to be added is constructed and solved to obtain the third weighted speech energy, as shown below: ; in, This represents the third weighted speech energy of the noise to be added.

5. The segmented energy-weighted noisy speech generation method as described in claim 4, characterized in that, The step of combining a preset target signal-to-noise ratio, the first weighted speech energy, and the third weighted speech energy to solve for the signal amplitude scaling factor in a preset signal-to-noise ratio calculation equation includes: The signal-to-noise ratio calculation equation is expressed as follows: ;in, Indicates the signal-to-noise ratio. This represents the first weighted speech energy. This represents the third-weighted speech energy; In the signal-to-noise ratio calculation equation, given the target signal-to-noise ratio, and substituting it into the expression for the third weighted speech energy, the signal amplitude scaling factor is obtained, as follows: .

6. The segmented energy-weighted noisy speech generation method as described in claim 5, characterized in that, The step of adjusting the original noise audio according to the signal amplitude scaling factor to generate noise audio to be added, and combining it with the original human voice audio to generate noisy speech that meets the target signal-to-noise ratio, includes: The signal relationship between the original noise audio and the noise to be added is expressed as follows: ;in, This represents the noise to be added, and m represents the signal amplitude scaling factor. Represents the original noise audio; Substituting the expression for the signal amplitude scaling factor, the generated audio with added noise is shown below: ; Based on the audio to be noised, and combined with the original human voice audio, noisy speech that meets the target signal-to-noise ratio is generated, as follows: ;in, This indicates noisy speech. This represents the original human voice audio.

7. A segmented energy-weighted noisy speech generation method as described in any one of claims 1-6, characterized in that, The size of the signal window is 10–30 ms.

8. A segmented energy-weighted noisy speech generation system, characterized in that, include: The system includes a data acquisition module, a weight calculation module, a weighted energy calculation module, a third weighted energy calculation module, a solution module, and a result output module. The data acquisition module is used to acquire original human voice audio and original noise audio of equal duration. The weight calculation module is used to calculate the energy weight of the audio signal of each signal window based on the preset signal window, respectively, on the original human voice audio and the original noise audio. The weighted energy calculation module is used to construct a weighted energy calculation equation based on the energy weight and sampling point energy of the audio signal in each signal window, and to calculate the first weighted speech energy of the original human voice audio and the second weighted speech energy of the original noise audio respectively. The third weighted energy calculation module is used to construct a third weighted energy calculation equation for the noise to be added based on a preset signal amplitude scaling factor between the original noise audio and the noise to be added, combined with the second weighted speech energy, and to calculate the third weighted speech energy of the noise to be added. The solution module is used to solve the signal amplitude scaling factor in a preset signal-to-noise ratio calculation equation by combining the preset target signal-to-noise ratio, the first weighted speech energy and the third weighted speech energy. The result output module is used to adjust the original noise audio according to the signal amplitude scaling factor, generate noise audio to be added, and combine it with the original human voice audio to generate noisy speech that meets the target signal-to-noise ratio.

9. A segmented energy-weighted noisy speech generation device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements the steps of a segmented energy-weighted noisy speech generation method according to any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of a segmented energy-weighted noisy speech generation method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Single-channel voice enhancement method and system

    CN102157156A

  • Combined suppression of noise and out-of-location signals

    CN103348408A