Human voice audio data processing method and system for reducing environmental noise

By employing a dynamic audio processing method based on emotion perception and noise source classification, the problems of noise suppression and emotion adaptation in multi-source environments are solved, achieving effective suppression of environmental noise and clear transmission of the singer's voice in occasions such as concerts.

CN121862162AActive Publication Date: 2026-04-14BEIJING STAR MECHANICAL & ELECTRICAL EQUIP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING STAR MECHANICAL & ELECTRICAL EQUIP
Filing Date
2026-01-20
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing audio processing technologies struggle to effectively suppress environmental noise in multi-source and dynamic environments while maintaining the clarity and realism of the singer's voice. Furthermore, they cannot adapt to emotional changes, resulting in a distorted performance atmosphere.

Method used

A dynamic audio processing method based on emotion perception and noise source classification is adopted. Audio signals are collected through a microphone array, multi-channel audio signal processing and sound source localization are performed, the main speech and environmental speech are identified, the emotional state is analyzed and the suppression intensity is adjusted to achieve dynamic adjustment.

Benefits of technology

It effectively suppresses environmental noise in complex acoustic scenarios, maintains the clarity and realism of the singer's voice, avoids frequent fluctuations in suppression intensity, and improves the stability and applicability of audio output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121862162A_ABST
    Figure CN121862162A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of quality inspection, and provides a human voice audio data processing method and system for reducing environmental noise. The method comprises the following steps of: acquiring an audio signal comprising a main voice emitting object and an environment voice emitting object group through a microphone array, and separating the audio signal based on multi-channel audio signal processing and sound source positioning to obtain a first audio signal corresponding to main voice and a second audio signal corresponding to environment voice; further analyzing the time-frequency characteristic of the second audio signal to obtain the emotion category and emotion intensity of the environment voice sending object group; and according to the change characteristics of the emotion intensity, adaptively determining and smoothly modulating the suppression intensity of the second audio signal, dynamically suppressing the environmental noise on the premise of ensuring the change continuity of the suppression intensity, and outputting the processed audio signal. According to the invention, adaptive modulation of environmental noise can be realized in a complex acoustic environment, and the method is suitable for scenes such as concerts and speeches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, specifically to a method and system for processing human voice audio data to reduce environmental noise. Background Technology

[0002] As the scale of large-scale events such as concerts and lectures continues to expand, the technical challenges faced by live sound systems are becoming increasingly complex. In particular, in multi-source environments, how to effectively suppress environmental noise (such as audience noise, background equipment noise, etc.) while maintaining the clarity and realism of the singer's (or speaker's) voice is an important issue in audio processing technology.

[0003] Existing audio processing technologies typically rely on static noise suppression algorithms. While these algorithms can reduce the impact of environmental noise to some extent, they often perform poorly in dynamic and complex environments. For example, in concert settings, the audience's emotional fluctuations cause the sound spectrum and intensity to change constantly. Traditional noise suppression methods cannot adapt to these rapid changes, easily leading to over-suppression or distortion of the singer's voice.

[0004] Furthermore, existing technologies rarely consider the needs of emotional perception and dynamic adjustment. In different emotional environments, simple noise suppression may excessively eliminate the audience's emotional expression (such as cheers), leading to a loss of the performance atmosphere. Therefore, how to maintain the atmosphere and audience response in concerts and other occasions while ensuring the singer's clear voice remains a pressing technical challenge that needs to be solved. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention proposes a dynamic audio processing method and system based on emotion perception and noise source classification. Through emotion recognition and dynamic separation of noise sources, it achieves efficient suppression of environmental noise while preserving the realism of the singer's or speaker's voice and the audience's emotional feedback.

[0006] This invention provides a method for processing human voice audio data to reduce environmental noise, comprising the following steps: S1, real-time acquisition of audio signals of speech emitting objects through a microphone array, wherein the speech emitting objects include at least one main speech emitting object and a group of environmental speech emitting objects; S2, through multi-channel audio signal processing and sound source localization, identifies and distinguishes the sound of the main speech source and the sound of the environmental speech source group, and performs spectral separation of the speech source to obtain the first audio signal and the second audio signal respectively; S3, based on the time-frequency characteristics of the second audio signal, analyze the emotional state of the group of people emitting environmental speech to obtain emotional state information, the emotional state information including emotion category and emotion intensity; S4, based on the emotional state information, smoothly enhance or smoothly weaken the suppression intensity of the second audio signal, and output the processed audio signal to a speaker or sound system.

[0007] The present invention also provides a human voice audio data processing system for reducing environmental noise, the system comprising: The audio acquisition module is used to acquire the audio signal of the speech emitting object in real time through a microphone array. The speech emitting object includes at least one main speech emitting object and a group of environmental speech emitting objects. The speech source separation module is used to identify and distinguish the sounds of the main speech source and the group of environmental speech sources through multi-channel audio signal processing and sound source localization, and to perform spectral separation of the speech sources to obtain the first audio signal and the second audio signal respectively. The emotion state analysis module is used to analyze the emotion state of the group of people emitting environmental speech based on the time-frequency characteristics of the second audio signal to obtain emotion state information, which includes emotion category and emotion intensity. The suppression intensity modulation module is used to smoothly enhance or weaken the suppression intensity of the second audio signal according to the emotional state information, and output the processed audio signal to a speaker or sound system.

[0008] The present invention also provides an electronic device comprising at least a processor and a memory, wherein the processor calls and executes program code in the memory to implement the steps of the method as described in any of the preceding claims.

[0009] The present invention also provides a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method as described in any of the preceding claims.

[0010] The present invention also provides a computer program product stored in a storage medium, the program product being executed by at least one processor to implement the steps of the method as described in any of the preceding claims.

[0011] Beneficial technical effects: This invention introduces a suppression intensity modulation mechanism based on the emotional state of environmental speech on the basis of multi-channel speech source separation. This makes the suppression process of environmental noise no longer dependent on fixed parameters or instantaneous judgment, but rather on smooth and adaptive adjustment according to the intensity of emotion and its changing patterns. This achieves stable suppression of environmental speech in complex and dynamic acoustic scenarios. This scheme can avoid sudden changes in sound quality caused by frequent fluctuations in suppression intensity, while ensuring the clarity and continuity of the speech of the main speech source, thus improving the stability and applicability of the overall audio output. Attached Figure Description

[0012] Figure 1 This is a flowchart illustrating a method for processing human voice audio data to reduce environmental noise, as disclosed in an embodiment of the present invention. Figure 2 This is a coordinate curve diagram of the separated first audio signal and second audio signal disclosed in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of the emotion classification model disclosed in an embodiment of the present invention; Figure 4 This is a schematic diagram of a human voice audio data processing system for reducing environmental noise, as disclosed in an embodiment of the present invention. Detailed Implementation

[0013] The following detailed description, in conjunction with specific embodiments, illustrates a method for processing human voice audio data to reduce environmental noise, as proposed in this invention. It should be noted that the following embodiments are merely illustrative of the technical solutions of this invention and are not intended to limit the scope of protection of this invention. Various equivalent modifications or substitutions made by those skilled in the art without departing from the concept of this invention should fall within the scope of protection of this invention.

[0014] The method proposed in this invention can be deployed in audio processing systems for performance venues, conference sound reinforcement systems, public address systems, or other terminals or servers with audio acquisition and processing capabilities, to stably output the main speech in scenarios with a large amount of environmental speech interference.

[0015] Please see Figure 1 This invention provides a method for processing human voice audio data to reduce environmental noise, comprising the following steps: S1, real-time acquisition of audio signals of speech emitting objects through a microphone array, wherein the speech emitting objects include at least one main speech emitting object and a group of environmental speech emitting objects; In practice, microphone arrays can be deployed at the front of the stage, around the podium, or at fixed locations within the sound system to simultaneously collect sound wave signals from multiple directions within a spatial range. The collected audio signals are multi-channel audio data containing mixed information from multiple sound sources, with each channel corresponding to the sampling results of different microphones in the array.

[0016] Among them, the main speech source object is used to represent the object that undertakes the main speech output function in the current scene, such as the singer in a concert, the speaker or host in a speech scene; the environmental speech source object group is used to represent the collection of multiple speech source objects that are not the main speech source and exist simultaneously in the same acoustic environment, such as the audience, the listeners or other background groups.

[0017] S2, through multi-channel audio signal processing and sound source localization, identifies and distinguishes the sound of the main speech source and the sound of the environmental speech source group, and performs spectral separation of the speech source to obtain the first audio signal and the second audio signal respectively; The aforementioned multi-channel audio signal is used as input, which contains mixed speech components from both the main speaker and the surrounding speech sources. To ensure that subsequent emotion analysis and suppression processing are performed only on the surrounding speech, it is necessary to first distinguish between different speech sources at the spatial and spectral levels.

[0018] As one embodiment, through multi-channel audio signal processing and sound source localization, the system identifies and distinguishes the sounds of the main speaker and the surrounding speech sources, and performs spectral separation of the speech sources to obtain a first audio signal and a second audio signal, respectively, including: S21, using time difference or phase difference methods, determine the spatial location of the main speech source and the group of environmental speech sources from the audio signal, and obtain their directional information; Specifically, the microphone array includes The multi-channel audio signals acquired by each microphone can be represented as follows: ; in, Indicates the first Each microphone at any time The collected audio signal.

[0019] Because the path lengths from the same sound source to different microphones differ, corresponding time or phase differences will occur in the time or frequency domain. Therefore, in this embodiment, the time difference is calculated by performing correlation analysis on the audio signals between any two microphone channels. ,For example: Alternatively, the phase difference can be calculated in the frequency domain: ,in Indicates to The result after frequency domain transformation.

[0020] Based on the aforementioned time difference or phase difference information, the spatial orientation of the corresponding sound source relative to the microphone array can be estimated, thereby obtaining the sound source orientation corresponding to the main speech source and the overall directional distribution information of the environmental speech source group. It is understood that this directional distribution information serves as a constraint condition for subsequent multi-channel audio signal differentiation processing.

[0021] S22, based on the directional information, a multi-channel audio signal processing method is used to separate the audio signal and distinguish the sound of the main speech source and the sound of the environmental speech source group; Specifically, based on the directional information obtained in step S21, a weighting process related to the sound source direction is applied to the multi-channel audio signal, enabling speech components from different spatial directions to be distinguished at the signal level. This process can be represented as: ; in, This represents a weighting coefficient related to the direction of the sound source. Indicates the direction of the object from which the main voice is emitted. This represents the set of directions corresponding to the group of objects emitting environmental speech.

[0022] Through the above processing, the speech components that are consistent with the direction of the main speech output are... The dominant component is the speech component, while speech components from the direction of the target group in the environment are extracted in a concentrated manner. This allows for the differentiation between the main speech and the ambient speech at the signal level.

[0023] S23, perform spectrum analysis on the separated audio signal to obtain the spectrum characteristics of the main speech source and the environmental speech source group, and generate the first audio signal and the second audio signal respectively.

[0024] Specifically, for and Perform spectral analysis processing separately, for example, by using short-time Fourier transform to convert the time-domain audio signal to a frequency-domain representation: ; ; in, Indicates frequency index, Indicates the time frame index.

[0025] Figure 2 The diagram shows the coordinate curves of the first audio signal (blue curve, Main voice) and the second audio signal (orange curve, Environment voice) after successful separation. The first audio signal exhibits a relatively periodic and continuous waveform, representing the speech components corresponding to the main speech source. The second audio signal exhibits a waveform with superimposed high-frequency fluctuations and random disturbances, representing the speech components corresponding to the environmental speech source group.

[0026] Through the above-described spectral analysis, spectral data representing the speech features of the main speech source and the speech features of the surrounding speech source group are obtained. In this embodiment, the signal containing the spectral features of the main speech source is used as the first audio signal, and the signal containing the spectral features of the surrounding speech source group is used as the second audio signal. The second audio signal serves as the input for subsequent emotion state analysis and suppression intensity modulation, while the first audio signal is used for joint output with the processed second audio signal.

[0027] S3, based on the time-frequency characteristics of the second audio signal, analyze the emotional state of the group of people emitting environmental speech to obtain emotional state information, the emotional state information including emotion category and emotion intensity; Since the second audio signal mainly contains speech components from the environmental speech emitters, its analysis can reflect the overall speech state changes of the environmental speech emitters in the current acoustic environment. To ensure that the emotion state analysis can adapt to the non-stationary characteristics of environmental speech, this step represents the emotion state of the environmental speech emitters layer by layer through three levels: time-frequency feature extraction, emotion category determination, and emotion intensity calculation.

[0028] As one embodiment, the emotional state information of the group of people emitting environmental speech is obtained by analyzing the time-frequency characteristics of the second audio signal, including: S31, perform time-domain and frequency-domain decomposition on the second audio signal to extract the time-frequency features of the environmental speech source group, including at least instantaneous frequency, amplitude spectrum and phase spectrum;

[0029] Specifically, let the second audio signal be represented in the time domain as follows: .

[0030] The second audio signal is segmented into frames according to a preset time window, and feature parameters characterizing speech changes are extracted in both the time and frequency domains within each time frame. Specifically: in the time domain, instantaneous frequency features reflecting changes in the rhythm of environmental speech are obtained through instantaneous characteristic analysis of the second audio signal; in the frequency domain, the corresponding amplitude and phase spectra are obtained through spectral analysis of the second audio signal, which are used to characterize the energy distribution and phase structure of different frequency components.

[0031] For example, in the frequency domain, the second audio signal can be represented as .in, Indicates frequency index, This represents the time frame index. Based on this frequency domain representation, the amplitude spectrum can be extracted. and phase spectrum Together with instantaneous frequency features, they constitute a set of time-frequency features used to describe the speech state of a group of objects emitting environmental speech.

[0032] S32, based on the time-frequency features, the emotion category of the target group of environmental speech is determined by the trained emotion classification model, and the emotion intensity information is calculated based on the changes of the determined emotion category in a continuous time period. Then, the emotion state is graded based on the emotion intensity information to obtain the emotion intensity corresponding to the emotion category.

[0033] The time-frequency features extracted in step S31 are used as input to a pre-trained emotion classification model. This model comprehensively determines the emotion category of the environmental speech source group within the current time period based on its integrated features in terms of rhythm, energy distribution, and spectral structure. Within each time frame or several consecutive time frames, the model outputs a corresponding emotion category identifier, representing the main emotion type of the environmental speech source group during that time period. This method allows for the acquisition of a time-varying sequence of emotion categories, which can then be used to calculate the subsequent emotion intensity.

[0034] For example: using the time-frequency features obtained by dividing the second audio signal into frames according to time windows as input, suppose a sample segment contains Each frame contains three types of features: amplitude spectrum. Phase spectrum Instantaneous frequency It is an instantaneous frequency sequence that can be obtained statistically by frequency band or full band.

[0035] Representing a segment as a three-way tensor / sequence: ,in For frequency points, This represents the number of frequency bands.

[0036] Please see Figure 3 The emotion classification model can be a multi-branch time-frequency feature fusion network, specifically including: Amplitude spectrum branch encoder: used to extract from... This extracts features of environmental speech energy distribution and rhythmic fluctuations. Its structure can employ 2D convolutional blocks (along frequency × time), with the output being an amplitude spectrum embedding vector. .

[0037] Phase spectrum branch encoder: used to extract from Extracting the time-varying pattern of the phase structure (phase information is sometimes more sensitive to group noise and choral ambient speech). The structure can be constructed using lightweight 2D convolutional blocks or 1D convolutions (along time), with the output being a phase spectrum embedding vector. .

[0038] Instantaneous frequency branch encoder: used to switch from Temporal features such as intonation variation, tremor level, and rhythm density are extracted. The structure can employ 1D convolution + gated recurrent unit (GRU) / temporal attention, and its output is an instantaneous frequency embedding vector. .

[0039] Feature fusion and classification head: Used to fuse the three embeddings and output the sentiment category. The fusion method uses interpretable weighted fusion, for example: .in, Based on a small fully connected layer Adaptive generation (i.e., attention weights). The classification head structure consists of a fully connected layer followed by a Softmax layer, outputting class probabilities. .

[0040] The final emotion category is .

[0041] After determining the emotion category, emotion intensity information is calculated based on the changes in the determined emotion category over a continuous time period, and the emotion state is further classified based on the emotion intensity information. Specifically, the emotion category determination results obtained over multiple consecutive time periods are constructed into a time series. .in, Indicates the first The emotion categories were determined within a specific time period. By quantifying the changes in the time series of these emotion categories, emotion intensity information reflecting the magnitude of changes in the emotional state of the environmental speech can be obtained. For example, when the emotion category changes frequently or significantly within adjacent time periods, the corresponding emotion intensity value is higher; when the emotion category remains relatively stable over a longer period, the corresponding emotion intensity value is lower.

[0042] S4, based on the emotional state information, smoothly enhance or smoothly weaken the suppression intensity of the second audio signal, and output the processed audio signal to a speaker or sound system.

[0043] Because the target group of environmental speech is usually group-oriented and asynchronous, its speech behavior often exhibits a combination of short-term fluctuations and periodic stability over time. Therefore, if the inhibition intensity is adjusted directly based on the emotional intensity value within a certain time period, it is easy to cause the inhibition intensity to change frequently over time, thereby affecting the continuity of audio output.

[0044] Based on the above considerations, this invention does not directly map emotional intensity to inhibition intensity. Instead, it introduces constraints such as the rate of change of emotional intensity, the allowable range of change, and time consistency to perform hierarchical control of the adjustment process of inhibition intensity.

[0045] For ease of explanation, let the emotion intensity obtained in the k-th time period be . The corresponding emotion category is .

[0046] Before modulating inhibition intensity, it is necessary to characterize the changes in emotional state over time. This involves differencing the emotional intensity over consecutive time periods to calculate the rate of change. Specifically, this involves differentiating the emotional intensity over adjacent time periods. The rate of change of emotion intensity between k and k can be expressed as: Furthermore, the emotion intensity sequence output in step S3 can be transformed into a rate of change sequence describing the magnitude of emotion changes. It should be noted that this rate of change is not directly used for suppression processing, but rather serves as an important reference for subsequently determining the scale of changes in suppression intensity.

[0047] The numerical range of emotion intensity may vary across different application scenarios or time periods. To ensure a consistent processing baseline for subsequent inhibition intensity modulation, the rate of change can be further normalized to bring it within a uniform scale range. ; in, This indicates the maximum rate of change observed within a preset time range.

[0048] As one embodiment, smoothly enhancing or weakening the suppression intensity of the second audio signal based on the emotional state information includes: S41, calculate the rate of change of emotional intensity based on the emotional state information, and dynamically adjust the target inhibition intensity through an adaptive filter; wherein, the target inhibition intensity is inversely proportional to the rate of change; Specifically, the normalized rate of change Next, the target inhibition strength is determined to match the characteristics of the current emotional change: the target inhibition strength is generated based on the rate of change using an adaptive filter. Their relationship can be represented as ,in This represents the baseline inhibition strength, used to define the overall range of inhibition strength values. It can be understood that, through the above mapping relationship, the determination of the target inhibition strength is related to the rate of emotional change, rather than directly tied to the intensity of emotion at a single moment.

[0049] It should be noted that the adaptive filter is used to dynamically modulate the suppression intensity. Its structure includes an input unit, a parameter update unit, and an output unit. The input unit receives the target suppression intensity or its candidate update value calculated from the emotional state information. The parameter update unit adaptively adjusts the filter's time constant and update coefficients according to the rate of change of emotional intensity and the trend of change of emotional category. The output unit generates the suppression intensity output for the current time period under the constraints of the time constant and update coefficients. During operation, the adaptive filter does not directly follow the instantaneous input change, but rather uses recursive calculations of historical suppression intensity states to make the current output gradually approach the target suppression intensity numerically, thus forming a controlled and continuous change process in the time dimension. This filtering process is also constrained by the allowable change range and the consistency judgment result, so that the filter's response speed matches the stability of the environmental speech emotional state.

[0050] As one embodiment, the step of dynamically adjusting the target suppression intensity through an adaptive filter includes: S411, determine the allowable change range of the target suppression intensity in adjacent time periods based on the change rate, and generate candidate change directions of the target suppression intensity based on the change rate; After obtaining the target suppression strength, a constraint on the magnitude of change is further introduced to limit the adjustment range of the suppression strength in adjacent time periods.

[0051] Specifically, the maximum permissible variation of the target suppression intensity within adjacent time periods is determined based on the rate of change: .

[0052] Based on this, the current target suppression intensity is compared with the suppression intensity of the previous time period. The comparisons are then used to generate candidate directions of change in inhibition intensity, which characterize the trend direction of inhibition intensity adjustment within the current time period. It is understood that these candidate directions are only used for subsequent consistency assessments and do not directly trigger inhibition intensity updates.

[0053] S412, in multiple consecutive time periods, the consistency of the candidate change direction is judged, and the target suppression intensity is updated by an adaptive filter only when the candidate change direction remains consistent within a preset time window; After obtaining candidate directions of change, the stability of the change trend over time is further considered. Therefore, consistency judgments are made on the candidate directions of change across multiple consecutive time periods.

[0054] Specifically, candidate change directions are statistically analyzed within a preset time window. The suppression strength is only allowed to be updated if the change directions remain consistent within this time window. In this case, the amount of suppression strength update is limited by the allowed change range. ; If the consistency condition is not met, the inhibitory intensity of the previous time period will remain unchanged, thereby avoiding frequent adjustments to the inhibitory intensity due to short-term emotional fluctuations.

[0055] S413, adjust the time constant of the adaptive filter according to the changing trend of the emotion category to control the update rate of the target inhibition intensity, and use the target inhibition intensity updated by the adaptive filter as the inhibition intensity of the second audio signal.

[0056] Even when the target inhibition strength can be updated, it is still necessary to further control the response speed of the update process on the time axis. To this end, this embodiment introduces a time constant as a dynamic parameter of the adaptive filter, and makes this time constant change with the changing trend of the emotion category, so that the update rate of the target inhibition strength is consistent with the stability of the emotion category.

[0057] In step S3, the emotion classification model outputs emotion category determination results over a continuous time period, thus obtaining the emotion category sequence over the continuous time period. .in, This indicates the length of the time window used to characterize the changing trends of emotion categories.

[0058] Based on this emotion category sequence, this embodiment statistically analyzes the changes in emotion categories within a window to obtain a trend quantity used to characterize the changing trend of emotion categories. For example, the number of times categories are switched within a window can be used to characterize trends of change: ; in, This is an indicator function that takes the value 1 when the condition within the parentheses is true, and 0 otherwise. Therefore, when... A larger value indicates more frequent switching of emotion categories within the window. A smaller value indicates that the emotion category is relatively stable within the window.

[0059] Based on the obtained trend volume Determine the time constant of the adaptive filter in the current time period. Time constant It is used to characterize the speed at which a filter responds to changes in the input, and its numerical value is used to control the speed of the target suppression intensity update process.

[0060] In this embodiment, the time constant can be set as a function of the trend quantity: ; in, and These represent the lower and upper limits of the allowed range of values ​​for the time constant, respectively. Therefore, when the emotion category is switched infrequently within the window, Closer When the number of switching operations is high, Closer .

[0061] To facilitate the application of time constants within discrete time periods, the time constants can be converted into filter update coefficients. ,For example: ; in, This represents the time interval between adjacent time periods.

[0062] In step S412, the candidate update amount or candidate target suppression strength under the condition of passing the consistency judgment has been obtained. In this step, by introducing... Rate control is applied to this update process so that the actual updated suppression strength does not directly equal the candidate value, but rather gradually approximates it according to a time constant. For example, if the consistency judgment allows updates, the target suppression strength can be updated as follows: ,in This represents the candidate suppression strength obtained under the constraint of allowable variation range. This represents the update result after introducing time constant control.

[0063] Therefore, when the emotion category is stable ( Smaller Larger (Closer to 1), the update process of inhibition intensity tends to maintain the state of the previous time period; when emotion categories change frequently ( Larger Smaller (Smaller), the updated results follow the candidate values ​​more closely.

[0064] After updating the time constant control as described above, perform step S42 to update the target suppression intensity. Further smoothing filtering is performed to obtain the final smoothing suppression strength used to suppress the second audio signal. .

[0065] S42, a smoothing filter is used to smooth the target suppression intensity, and the smoothed target suppression intensity is applied to the second audio signal.

[0066] The suppression strength obtained after step S413 The following conditions have been met: 1) its update only occurs when the consistency judgment in step S412 passes; 2) its update rate is constrained by the time constant, thus avoiding excessively rapid numerical jumps. However, since the update of the suppression intensity is still based on discrete time intervals, numerical changes may still occur between adjacent time intervals. Therefore, it is necessary to further smooth the suppression intensity in time before directly applying it to the audio signal. Specifically:

[0067] When entering the smoothing process, the current time period has been obtained. Internally determined inhibition strength And the smoothing suppression intensity that has been applied to the audio signal in the previous time period. This step processes the inhibition intensity values ​​within the two consecutive time periods mentioned above. Its purpose is to further reduce numerical abrupt changes between adjacent time periods while maintaining the overall trend of inhibition intensity changes.

[0068] In this embodiment, the suppression intensity is smoothed over time using a smoothing filter. Specifically, the suppression intensity is smoothed. It can be calculated as follows: ,in The smoothing coefficient controls the weight ratio of the suppression intensity of the current time period to the smoothing result of the previous time period in the final output.

[0069] The above calculation method makes the smoothing suppression strength exhibit continuous variation characteristics in the time dimension, that is, when Compared to When changes occur, Instead of immediately and completely following the change, it gradually transitions to the new inhibition intensity value in a controlled manner.

[0070] It should be noted that the smoothing process does not change the already determined direction and magnitude constraints of the suppression intensity change, but rather further processes the temporal continuity of the suppression intensity to make it more suitable for direct application to audio signals.

[0071] After completing the smoothing calculation of the suppression intensity, the smoothed suppression intensity is obtained. This is used as a control parameter for suppressing the second audio signal within the current time period. In practice, the smoothing suppression intensity remains constant within the current time period and participates in the new smoothing calculation process as a historical state when the next time period arrives. In this way, the suppression intensity forms a continuous and controllable change curve over a continuous time period, rather than a discrete and abrupt control sequence.

[0072] After obtaining the smoothing suppression strength, it is applied to the second audio signal to suppress the speech components corresponding to the target group of environmental speech. Specifically, let the second audio signal be represented as follows within the current time period: The signal after suppression can be expressed as In this way, the suppression intensity is applied to the second audio signal in a continuous and controlled manner, so that the speech components of the environmental speech source group are smoothly weakened in the time dimension.

[0073] Please see Figure 4 The present invention provides a human voice audio data processing system 100 for reducing environmental noise, the system comprising: The audio acquisition module 11 is used to acquire the audio signal of the voice emitting object in real time through a microphone array. The voice emitting object includes at least one main voice emitting object and a group of environmental voice emitting objects. The speech source separation module 12 is used to identify and distinguish the sounds of the main speech source and the environmental speech source group through multi-channel audio signal processing and sound source localization, and to perform spectral separation of the speech source to obtain the first audio signal and the second audio signal respectively. The emotion state analysis module 13 is used to analyze the emotion state of the group of people emitting environmental speech based on the time-frequency characteristics of the second audio signal to obtain emotion state information, which includes emotion category and emotion intensity. The suppression intensity modulation module 14 is used to smoothly enhance or weaken the suppression intensity of the second audio signal according to the emotional state information, and output the processed audio signal to a speaker or sound system.

[0074] As one embodiment, the voice source separation module 12 is configured as follows: The spatial positions of the main speech source and the group of environmental speech sources are determined from the audio signal using time difference or phase difference methods, and their directional information is obtained. Based on the directional information, a multi-channel audio signal processing method is used to separate the audio signal and distinguish the sound of the main speech source from the sound of the environmental speech source group. Spectral analysis is performed on the separated audio signals to obtain the spectral characteristics of the main speech source and the environmental speech source group, and the first audio signal and the second audio signal are generated respectively.

[0075] As one embodiment, the emotion state analysis module 13 is configured to: The second audio signal is decomposed in the time domain and frequency domain to extract the time-frequency features of the environmental speech source group, including at least the instantaneous frequency, amplitude spectrum and phase spectrum; Based on the time-frequency features, the emotion category of the target group of environmental speech is determined by a trained emotion classification model. The emotion intensity information is calculated based on the changes of the determined emotion category over a continuous time period. Then, the emotion state is graded based on the emotion intensity information to obtain the emotion intensity corresponding to the emotion category.

[0076] As one embodiment, the suppression intensity modulation module 14 is configured to: The rate of change of emotional intensity is calculated based on the emotional state information, and the target inhibition intensity is obtained by dynamically adjusting it through an adaptive filter; wherein, the target inhibition intensity is inversely proportional to the rate of change. The target suppression intensity is smoothed using a smoothing filter, and the smoothed target suppression intensity is applied to the second audio signal.

[0077] As one embodiment, the suppression intensity modulation module 14 is configured to: The allowable range of change of the target suppression intensity within adjacent time periods is determined based on the rate of change, and candidate directions of change of the target suppression intensity are generated based on the rate of change. Consistency judgment is performed on the candidate change direction over multiple consecutive time periods. The target suppression intensity is updated only when the candidate change direction remains consistent within a preset time window. The time constant of the adaptive filter is adjusted according to the changing trend of the emotion category to control the update rate of the target inhibition intensity, and the target inhibition intensity updated by the adaptive filter is used as the inhibition intensity of the second audio signal.

[0078] This invention also provides an electronic device, which includes at least a processor and a memory, wherein the processor calls and executes program code in the memory to implement the steps of the method as described in any of the preceding claims.

[0079] This invention also provides a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the method as described in any of the preceding claims.

[0080] This invention also provides a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the method as described in any of the preceding claims.

[0081] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.

[0082] Those skilled in the art will understand that the modules in the system of the embodiments can be distributed in the system of the embodiments as described in the embodiments, or they can be located in one or more systems different from this embodiment with corresponding changes. The modules of the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing human voice audio data to reduce environmental noise, characterized in that, The methods and steps include the following: S1, real-time acquisition of audio signals of speech emitting objects through a microphone array, wherein the speech emitting objects include at least one main speech emitting object and a group of environmental speech emitting objects; S2, through multi-channel audio signal processing and sound source localization, identifies and distinguishes the sound of the main speech source and the sound of the environmental speech source group, and performs spectral separation of the speech source to obtain the first audio signal and the second audio signal respectively; S3, based on the time-frequency characteristics of the second audio signal, analyze the emotional state of the group of people emitting environmental speech to obtain emotional state information, the emotional state information including emotion category and emotion intensity; S4, based on the emotional state information, smoothly enhance or smoothly weaken the suppression intensity of the second audio signal, and output the processed audio signal to a speaker or sound system.

2. The method for processing human voice audio data to reduce environmental noise according to claim 1, characterized in that: Through multi-channel audio signal processing and sound source localization, the system identifies and distinguishes between the main speaker and the surrounding speech sources, and performs spectral separation of the speech sources to obtain a first audio signal and a second audio signal, including: S21, using time difference or phase difference methods, determine the spatial location of the main speech source and the group of environmental speech sources from the audio signal, and obtain their directional information; S22, based on the directional information, a multi-channel audio signal processing method is used to separate the audio signal and distinguish the sound of the main speech source and the sound of the environmental speech source group; S23, perform spectrum analysis on the separated audio signal to obtain the spectrum characteristics of the main speech source and the environmental speech source group, and generate the first audio signal and the second audio signal respectively.

3. The method for processing human voice audio data to reduce environmental noise according to claim 2, characterized in that: Based on the time-frequency characteristics of the second audio signal, the emotional state of the target group emitting the environmental speech is analyzed to obtain emotional state information, including: S31, perform time-domain and frequency-domain decomposition on the second audio signal to extract the time-frequency features of the environmental speech source group, including at least instantaneous frequency, amplitude spectrum and phase spectrum; S32, based on the time-frequency features, the emotion category of the target group of environmental speech is determined by the trained emotion classification model, and the emotion intensity information is calculated based on the changes of the determined emotion category in a continuous time period. Then, the emotion state is graded based on the emotion intensity information to obtain the emotion intensity corresponding to the emotion category.

4. A method for processing human voice audio data to reduce environmental noise according to claim 3, characterized in that: The suppression intensity of the second audio signal is smoothly enhanced or weakened based on the emotional state information, including: S41, calculate the rate of change of emotional intensity based on the emotional state information, and dynamically adjust the target inhibition intensity through an adaptive filter; wherein, the target inhibition intensity is inversely proportional to the rate of change; S42, a smoothing filter is used to smooth the target suppression intensity, and the smoothed target suppression intensity is applied to the second audio signal.

5. A method for processing human voice audio data to reduce environmental noise according to claim 4, characterized in that: The target suppression strength is obtained by dynamically adjusting the adaptive filter, including: S411, determine the allowable change range of the target suppression intensity in adjacent time periods based on the change rate, and generate candidate change directions of the target suppression intensity based on the change rate; S412, in multiple consecutive time periods, the consistency of the candidate change direction is judged, and the target suppression intensity is updated by an adaptive filter only when the candidate change direction remains consistent within a preset time window; S413, adjust the time constant of the adaptive filter according to the changing trend of the emotion category to control the update rate of the target inhibition intensity, and use the target inhibition intensity updated by the adaptive filter as the inhibition intensity of the second audio signal.

6. A human voice audio data processing system for reducing environmental noise, characterized in that, The system includes: The audio acquisition module is used to acquire the audio signal of the speech emitting object in real time through a microphone array. The speech emitting object includes at least one main speech emitting object and a group of environmental speech emitting objects. The speech source separation module is used to identify and distinguish the sounds of the main speech source and the group of environmental speech sources through multi-channel audio signal processing and sound source localization, and to perform spectral separation of the speech sources to obtain the first audio signal and the second audio signal respectively. The emotion state analysis module is used to analyze the emotion state of the group of people emitting environmental speech based on the time-frequency characteristics of the second audio signal to obtain emotion state information, which includes emotion category and emotion intensity. The suppression intensity modulation module is used to smoothly enhance or weaken the suppression intensity of the second audio signal according to the emotional state information, and output the processed audio signal to a speaker or sound system.

7. A human voice audio data processing system for reducing environmental noise according to claim 6, characterized in that: The speech source separation module is configured as follows: The spatial positions of the main speech source and the group of environmental speech sources are determined from the audio signal using time difference or phase difference methods, and their directional information is obtained. Based on the directional information, a multi-channel audio signal processing method is used to separate the audio signal and distinguish the sound of the main speech source from the sound of the environmental speech source group. Spectral analysis is performed on the separated audio signals to obtain the spectral characteristics of the main speech source and the environmental speech source group, and the first audio signal and the second audio signal are generated respectively.

8. An electronic device, characterized in that, The electronic device includes at least a processor and a memory, wherein the processor calls and executes program code in the memory to implement the steps of the method as described in any one of claims 1 to 5.

9. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 5.

10. A computer program product, characterized in that, The program product is stored in a storage medium and is executed by at least one processor to implement the steps of the method as claimed in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice affect modification

    CN106992013A

  • Speech translation interaction method and system

    CN108228575A

  • Hearing aid audio optimization method and system based on intelligent scene recognition

    CN120111426A

  • Noise reduction regulation and control method for earphone and noise reduction earphone

    CN120833774A

  • Systems and methods for automatic-generation of soundtracks for live speech audio

    EP3276617A1