Audio data processing method and apparatus, medium, device, and program product

By generating simulated noisy data and performing first-order and second-order speech enhancement, the problem of insufficient data in AI speech processing is solved, diversified audio data synthesis is achieved, acquisition costs are reduced, and data processing efficiency is improved.

CN119993185BActive Publication Date: 2025-10-14TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510258693.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-01
Publication Date
2025-10-14
Estimated Expiration
2041-12-01

AI Technical Summary

Technical Problem

Existing technologies lack sufficient quantity or type of training data in AI speech processing, resulting in overfitting and poor recognition effects, and manual collection of audio data consumes a lot of manpower and financial resources.

Method used

By acquiring pure speech audio data and noisy audio data, generating simulated noisy data, and simulating the changes in audio after passing through space, performing first-order and second-order speech enhancement operations, and synthesizing diverse target audio data.

Benefits of technology

It effectively improves the diversity of simulated audio data sets, reduces the need for manual collection, reduces costs, and improves data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993185B_ABST
    Figure CN119993185B_ABST
Patent Text Reader

Abstract

The application discloses an audio data processing method and device, a medium, equipment and a program product, which are applied to AI noise reduction, AI echo cancellation and the like in artificial intelligence (AI) and machine learning. The method comprises the following steps: acquiring original audio data collected, wherein the original audio data comprises pure speech audio data and noise audio data; generating simulation noise data according to the pure speech audio data and the noise audio data in the original audio data; generating target audio data used for simulating changes of audio after the audio is transmitted through a space according to the original audio data or the simulation noise data; and performing a speech enhancement operation to obtain enhanced target audio data, wherein the speech enhancement operation comprises performing first-order speech enhancement on the original audio data, regenerating updated target audio data based on the data of the first-order speech enhancement, and performing second-order speech enhancement on the target audio data and / or the updated target audio data, so as to improve diversity of a simulation audio data set.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese patent application No. 202111456334.6, titled "Audio data processing method and device, medium, equipment and program product", filed on December 1, 2021 with the China Patent Office. TECHNICAL FIELD

[0002] The present application relates to the technical field of speech signal processing, in particular to an audio data processing method and device, medium, equipment and program product. BACKGROUND

[0003] With the continuous development of speech signal processing technology and artificial intelligence (AI) technology, there are more and more tasks of processing speech through AI, such as AI speech noise reduction, echo cancellation, etc. In the processing process, a large amount of various noise audio and various special scene audio such as echo audio, which can be used for training, need to be collected, and such collection often consumes a lot of manpower and financial resources. In AI speech processing, if there is a lack of sufficient amount or type of training data, it is easy to cause overfitting, poor recognition effect and other problems. SUMMARY

[0004] The embodiments of the present application provide an audio data processing method, device, medium, equipment and program product, which improve the diversity of the simulation audio data synthesis method.

[0005] In one aspect, an audio data processing method is provided, the method comprising:

[0006] obtaining collected original audio data, the original audio data comprising pure speech audio data and noise audio data;

[0007] generating simulation noisy data according to the pure speech audio data and the noise audio data in the original audio data;

[0008] generating target audio data for simulating changes of audio after spatial transmission according to the original audio data or the simulation noisy data, wherein the spatial transmission comprises transmission through a loudspeaker or transmission through a near-end loudspeaker and a microphone;

[0009] performing speech enhancement operation to obtain enhanced target audio data, wherein the speech enhancement comprises first-order speech enhancement and second-order speech enhancement, and specifically:

[0010] performing the first-order speech enhancement operation on the original audio data before inputting the target audio data into the speech model, to obtain first-order speech enhanced original audio data, the first-order speech enhanced original audio data being used to generate updated target audio data; the first-order speech enhancement including audio speed change, volume adjustment, random displacement, noise enhancement, and multiplication enhancement;

[0011] performing random information loss processing on the target audio data and / or the updated target audio data in a feature dimension in a time-frequency domain, to obtain second-order enhanced target audio data.

[0012] In another aspect, an audio data processing apparatus is provided, the apparatus comprising:

[0013] an obtaining unit configured to obtain collected original audio data, the original audio data including clean speech audio data and noise audio data;

[0014] a generating unit configured to generate simulated noise data according to the clean speech audio data and the noise audio data in the original audio data; and

[0015] generate target audio data for simulating changes of audio after spatial transmission, according to the original audio data or the simulated noise data, wherein the spatial transmission includes transmission through a loudspeaker or transmission through a near-end loudspeaker and a microphone.

[0016] an enhancing unit configured to perform speech enhancement to obtain enhanced target audio data, wherein the speech enhancement includes first-order speech enhancement and second-order speech enhancement, and the enhancing unit is configured to:

[0017] perform the first-order speech enhancement operation on the original audio data before inputting the target audio data into the speech model, to obtain first-order speech enhanced original audio data, the first-order speech enhanced original audio data being used to generate updated target audio data; the first-order speech enhancement including audio speed change, volume adjustment, random displacement, noise enhancement, and multiplication enhancement;

[0018] perform random information loss processing on the target audio data and / or the updated target audio data in a feature dimension in a time-frequency domain, to obtain second-order enhanced target audio data.

[0019] In another aspect, a computer readable storage medium is provided, the computer readable storage medium storing a computer program, the computer program being adapted to be loaded by a processor to perform steps in the audio data processing method according to any one of the above embodiments.

[0020] In another aspect, a computer device is provided, which includes a processor and a memory, the memory storing a computer program, the processor being configured to perform the steps of the audio data processing method of any of the above embodiments by invoking the computer program stored in the memory.

[0021] In another aspect, a computer program product is provided, which includes computer instructions for implementing the steps of the audio data processing method of any of the above embodiments when executed by a processor.

[0022] The embodiments of the present application obtain the collected original audio data, the original audio data including clean speech audio data and noise audio data; generate simulated noisy data according to the clean speech audio data and the noise audio data in the original audio data; generate target audio data for simulating changes of audio after spatial transmission according to the original audio data or the simulated noisy data, wherein the spatial transmission includes transmission through a loudspeaker or transmission through a near-end loudspeaker and a microphone; perform speech enhancement to obtain enhanced target audio data, wherein the speech enhancement includes first-order speech enhancement and second-order speech enhancement, specifically: performing first-order speech enhancement on the original audio data before inputting the target audio data into a speech model to obtain first-order speech enhanced original audio data, the first-order speech enhanced original audio data being used to generate updated target audio data; the first-order speech enhancement includes audio speed change, volume adjustment, random displacement, noise enhancement, and multiplication enhancement; and performing random information loss processing on the target audio data and / or the updated target audio data in the feature dimension in the time-frequency domain inside the speech model to obtain second-order enhanced target audio data. A large amount of clean human voice audio and various types of noise audio are used to describe changes in the speech space propagation path through mathematical language, and various simulated target audio data are synthesized. Compared with the large amount of manpower and material resources consumed by manual collection of audio data in the related art, the original audio data is used for audio data processing, the changes of audio transmission through various spaces are simulated through mathematical language, and diversified target audio data is automatically generated in batches, and a more complete simulated audio data synthesis method is proposed. In addition, speech enhancement is performed on the generated target audio data, the speech enhancement includes first-order speech enhancement on the original audio data, and the updated target audio data is regenerated based on the first-order speech enhanced data, and second-order speech enhancement is performed on the target audio data and / or the updated target audio data, thereby further improving the diversity of the simulated audio data set. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical methods in the embodiments of the present application, the drawings needed in the description of the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0024] Figure 1 A flowchart of an audio data processing method provided by an embodiment of the present application is shown.

[0025] Figure 2 Another flowchart of an audio data processing method provided by an embodiment of the present application is shown.

[0026] Figure 3 An example diagram of an audio data processing method provided by an embodiment of the present application is shown.

[0027] Figure 4 Another example diagram of an audio data processing method provided by an embodiment of the present application is shown.

[0028] Figure 5 Another example diagram of an audio data processing method provided by an embodiment of the present application is shown.

[0029] Figure 6 Another example diagram of an audio data processing method provided by an embodiment of the present application is shown.

[0030] Figure 7 A structural diagram of an audio data processing device provided by an embodiment of the present application is shown.

[0031] Figure 8 A structural diagram of an audio data processing device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0032] The technical methods in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0033] The embodiments of the present application provide an audio data processing method, an audio data processing device, a medium and equipment. Specifically, the method of the embodiments of the present application can be executed by a computer device, which can be a terminal or a server or the like. The embodiments of the present application can be applied to the research of artificial intelligence AI noise reduction and artificial intelligence AI echo cancellation in artificial intelligence and machine learning, and can also be used as an auxiliary method for improving data diversity in the research process of voice recognition and speaker recognition and the like.

[0034] First, some of the nouns or terms that appear in the description of the embodiments of the present application are explained as follows:

[0035] Artificial Intelligence (AI): is to use digital computer or digital computer controlled machine simulation, extension and expansion of human intelligence, perception of environment, acquisition of knowledge and use of knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is the design principle and implementation method of various intelligent machines, so that the machine has the functions of perception, reasoning and decision making. The AI voice model in the present application is to process the voice audio through the AI machine learning model to obtain the corresponding analysis result, such as AI voice noise reduction and echo cancellation.

[0036] Signal-to-noise ratio (SNR): represents the ratio of the amplitudes of useful signal and noise signal.

[0037] Perceptual Evaluation of Speech Quality (PESQ): Perceptual Evaluation of Speech Quality is an objective and full-reference speech quality evaluation method. Its algorithm needs a noisy attenuated signal and an original reference signal, which can provide a subjective prediction value for objective speech quality evaluation, and the score is between -0.5 and 4.5. The higher the score, the better the speech quality.

[0038] Blockchain system: can be a distributed system formed by connecting clients, multiple nodes (any form of computing devices such as servers and user terminals in the network) through network communication. The nodes form a peer-to-peer (P2P, Peer To Peer) network, and the P2P protocol is an application layer protocol running on the Transmission Control Protocol (TCP, Transmission Control Protocol) protocol. In a distributed system, any machine such as a server or terminal can join to become a node, and the node includes a hardware layer, an intermediate layer, an operating system layer and an application layer.

[0039] The rise of data-driven AI algorithms and the wide range of productized applications have brought great convenience to our lives. Data-driven AI algorithms are to enable models to understand data and learn knowledge from historical data to predict unknown data. Therefore, to produce AI models with strong generalization ability, machines need to be experienced and knowledgeable, and accumulate enough rich data.

[0040] Among them, the audio database required for training the AI voice model needs high cost in practice. For example, in the training of AI voice model for processing audio such as AI voice recognition, or in actual application, a large amount of audio data available for training is required. At present, there is a general basic audio database, but the diversity of the general data is not enough to adapt to various business needs, and a large amount of various noise audio, echo audio, and audio data for various special application scenarios need to be collected manually. However, such collection requires a large amount of manpower, financial resources and material resources.

[0041] In another type of technology, an audio data augmentation method for speech recognition is proposed. The augmentation strategy includes using function sparse image warping to apply time warping, randomly selecting audio frequency domain channels for shielding, and randomly selecting audio time domain channels for shielding. However, more diverse data sets are still needed in some special application scenarios.

[0042] From the beginning of AI voice technology to the development so far, a complete industry chain including upstream, midstream and downstream has been formed. The downstream industry application of intelligent voice technology is diversified, and one-stop service is widely demanded. The consumer application field currently includes but is not limited to: chat APP, smart hardware, smart home, vehicle-mounted system, etc. The present application proposes an audio data processing method for the technical field of AI voice model training and other needs of various types of audio data, which can be applied to the research of artificial intelligence AI noise reduction, artificial intelligence AI echo cancellation technology in machine learning, and can also be used as an auxiliary method to improve data diversity in the process of speech recognition, speaker recognition and other technical research, etc. A more complete audio data processing method is proposed, which effectively improves the diversity of the audio data set.

[0043] In order to better understand the technical method provided by the embodiments of the present application, the application scenarios to which the technical method provided by the embodiments of the present application are applied will be briefly introduced below. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application and are not limited. For example, the audio data processing method is executed by a computer device, wherein the computer device can be a terminal device or a server.

[0044] The embodiments of the present application can be implemented in combination with cloud technology or blockchain network technology. For example, the audio data processing method disclosed in the embodiments of the present application, wherein the data can be saved on the blockchain, for example: original audio data, pure voice audio data, noise audio data, simulated noise data, target audio data, and enhanced target audio data, can be saved on the blockchain.

[0045] To facilitate the storage and query of the original audio data, the clean speech audio data, the noise audio data, the simulated noise data, the target audio data, and the enhanced target audio data, optionally, the audio data processing method further includes: sending the original audio data, the clean speech audio data, the noise audio data, the simulated noise data, the target audio data, and the enhanced target audio data to a blockchain network, so that a node of the blockchain network fills the original audio data, the clean speech audio data, the noise audio data, the simulated noise data, the target audio data, and the enhanced target audio data to a new block, and when a consensus is reached on the new block, the new block is appended to the tail of the blockchain. The embodiments of the application can store the original audio data, the clean speech audio data, the noise audio data, the simulated noise data, the target audio data, and the enhanced target audio data on the chain, realize backup of the record, and when the target audio data and the enhanced target audio data are needed, the corresponding target audio data and the enhanced target audio data can be directly and quickly obtained from the blockchain, thereby improving the efficiency of audio data processing.

[0046] The following will be described in detail. It should be noted that the order of description of the following embodiments is not limited as the priority order of the embodiments.

[0047] The embodiments of the application provide an audio data processing method, and the embodiments of the application take the audio data processing by a computer device as an example for description.

[0048] Please refer to Figure 1 , Figure 1 The flowchart of the audio data processing method provided by the embodiments of the application is shown in the figure, and the method includes:

[0049] In step 110, the original audio data collected is obtained, and the original audio data includes clean speech audio data and noise audio data.

[0050] Specifically, the original audio data at least includes clean speech audio data and noise audio data, and can further include room impulse response.

[0051] The clean speech audio data includes clean human voice audio data. The data category is not limited, and can include multi-national language, multi-dialect, and human voice humming music. The clean human voice audio data can be collected in advance by various ways, including obtaining from a public audio database, a special audio database, or manual collection. For example, a large amount of clean human voice audio is recorded from a real scene.

[0052] Noise audio data includes various types of noise, such as pure subway noise, traffic noise, natural background noise, vehicle noise, indoor and outdoor noise, and other common noise scenarios. Noise audio data can be collected in advance through various methods, including public audio databases, proprietary audio databases, or manual collection.

[0053] Step 120 : Generate simulated noisy data based on the clean speech audio data and the noisy audio data in the original audio data.

[0054] Specifically, after obtaining pure voice audio data and noisy audio data. For example, the clean human voice audio data can be added to the noisy audio data to synthesize a new simulated noisy data, and then a large amount of simulated noisy data can be synthesized by mixing the pure voice audio data with multiple types of noisy audio data. Among them, the simulated noisy data contains a voice signal and its background noise signal. In natural life, a large amount of audio data includes noisy audio data of voice signals and background noise signals. In practical applications of audio, such as the application of AI echo cancellation models, the input of the far-end microphone is human voice and background noise, and the input of the near-end microphone is also human voice and background noise. The input of the far-end microphone and the near-end microphone can be simulated by synthesizing simulated noisy data.

[0055] The method for synthesizing aliased pure speech audio data and various types of noisy audio data can be implemented according to the signal-to-noise ratio (SNR), and the noisy data can be synthesized according to the audio signal-to-noise ratio, which is the ratio of the normal sound signal strength to the noise signal strength. For example:

[0056] The signal-to-noise ratio calculation method can be expressed as formula (1):

[0057]

[0058] Where SNR represents the signal-to-noise ratio in dB, s(t) represents pure speech audio data, ∑ t s 2 (t) represents the speech energy of pure speech audio data, n(t) represents the noisy audio data, ∑ t n 2 (t) represents the noise energy of the noisy audio data.

[0059] When synthesizing pure speech audio data and multiple types of noise audio data, the noise energy can be adjusted to α times the original value, that is, αn(t). The signal-to-noise ratio is expressed as formula (2):

[0060]

[0061] Where q represents the signal-to-noise ratio (SNR) under the adjusted noise energy ratio.

[0062] Thus, the calculation formula of the noise energy adjustment ratio a can be represented as formula (3):

[0063]

[0064] In the synthesis of pure speech audio data and various types of noise audio data, by giving a preset signal-to-noise ratio SNR, that is, q in formula (3), optionally, the signal-to-noise ratio q can be randomly selected as an integer in the range of (-5, 20). Then the noise energy adjustment ratio a can be obtained by formula (3), and then the pure speech audio data s(t) and the noise audio data n(t) are synthesized according to the signal-to-noise ratio to generate the simulated noisy data, which can be represented as formula (4):

[0065] mix(t) = s(t) + an(t) (4).

[0066] Optionally, according to the pure speech audio data and the noise audio data in the original audio data, the step of generating the simulated noisy data further comprises:

[0067] The multiplicative noise audio data in the noise audio data is converted into additive noise audio data through homomorphic filtering processing;

[0068] According to the signal-to-noise ratio, the pure speech audio data and the additive noise audio data are synthesized to obtain the simulated noisy data.

[0069] Specifically, the collected noise audio data can be additive noise audio data or multiplicative noise audio data. The additive noise audio data and the multiplicative noise audio data are two widely used noise types. The additive noise audio data includes thermal noise, shot noise, etc. The relationship between the additive noise audio data and the signal is addition, and the additive noise audio data exists regardless of whether there is a signal. The multiplicative noise audio data is generally caused by an imperfect channel, and the relationship between the multiplicative noise audio data and the signal is multiplication, and the signal and the multiplicative noise audio data exist at the same time. The additive noise audio data can be used to simulate background noise, and the multiplicative noise audio data can be used to simulate the time-varying or nonlinearity of the system.

[0070] It can be understood that the additive noise audio data in the interference of noise to speech is that the signals are added in the time domain, or from the energy angle, the additive noise audio data and the speech are in the superposition relationship of sound intensity, and the two form a noisy speech signal together by the microphone.

[0071] The multiplicative noise audio data refers to a convolution relationship between noise and speech in the time domain and a multiplication relationship between noise and speech in the frequency domain. The multiplicative noise audio data can be converted into additive noise audio data through transformation. For example, the multiplicative noise audio data or the convolution noise audio data can be converted into additive noise audio data through homomorphic filtering.

[0072] The conversion of the multiplicative noise audio data into additive noise audio data can include the following steps:

[0073] Firstly, the multiplicative noise audio data can be expressed as:

[0074] x(t)=x1(t)*x2(t) (5);

[0075] wherein x(t) represents the multiplicative noise audio data, x1(t) represents speech in the multiplicative noise audio data, and x2(t) represents noise in the multiplicative noise audio data.

[0076] The Z transformation is performed on the formula (5) to convert the convolution signal into a product signal, as shown in the formula (6):

[0077] Z[x(t)]=X(z)=X1(z)*X2(z) (6)。

[0078] Then, the logarithmic operation is performed on both sides of the formula (6) to convert the multiplication operation into an addition operation, as shown in the formula (7):

[0079]

[0080] Then, the inverse Z transformation is performed on to convert the logarithmic z-domain signal into a time-domain signal, as shown in the formula (8):

[0081]

[0082] In this way, the multiplicative noise audio data x(n) is converted into additive noise audio data

[0083] Further, the pure speech audio data and the additive noise audio data are synthesized according to the signal-to-noise ratio to obtain the simulation noise data.

[0084] The specific synthesis method can be synthesized through the above formulas (1)-(4), that is, after a preset signal-to-noise ratio SNR is given, the specific simulation noise data is synthesized through the above formulas (1)-(4).

[0085] In step 130, the target audio data used to simulate the change of the audio after the spatial transmission is generated according to the original audio data or the simulation noise data.

[0086] The spatial transmission includes transmission through a speaker or transmission through a near-end speaker and a microphone.

[0087] Specifically, the change in the spatial transmission of the audio data can be described using mathematical language. The change in the spatial transmission can include a change in transmission through a speaker or a change in transmission through a near-end speaker and a microphone. Accordingly, the target audio data can include reverberation audio data or sound back audio data, etc.

[0088] Optionally, the target audio data includes reverberation audio data, and the step 130 includes:

[0089] Step 1301: generating, according to at least one of the simulated noise data, the clean speech audio data, and the noise audio data, simulated speaker audio for simulating a change in audio after the audio passes through a speaker;

[0090] Step 1302: generating, according to the simulated speaker audio and a room impulse response, reverberation audio data.

[0091] Optionally, the step 1301 includes:

[0092] processing at least one of the simulated noise data, the clean speech audio data, and the noise audio data as a speaker input signal to obtain an audio signal maximum value;

[0093] generating, according to the audio signal maximum value and the speaker input signal, speaker power amplifier audio for simulating a change in audio after the audio passes through a power amplifier saturation region inside a speaker;

[0094] performing first non-linear conversion on the speaker power amplifier audio to obtain non-linear speaker power amplifier audio;

[0095] processing the non-linear speaker power amplifier audio using a non-linear function to generate the simulated speaker audio.

[0096] Specifically, at least one of the simulated noise data, the clean speech audio data, and the noise audio data is processed as a speaker input signal x(t) to obtain an audio signal maximum value x max . Optionally, x max may also be set as a proportion of the input signal maximum value, such as 80% of the maximum value.

[0097] Further, the speaker power amplifier audio for simulating a change in audio after the audio passes through a power amplifier saturation region inside a speaker is generated according to the audio signal maximum value and the speaker input signal. Specifically, the speaker input signal can be the simulated noise data, the clean speech audio data, or the noise audio data. For example, according to the simulated noise data, the simulated noise data can be synthesized through the above formulas (1)-(8), and the simulated noise data can be represented as

[0098] x(t) = s(t) + an(t) (9);

[0099] where s(t) represents the clean speech audio data, n(t) represents the noise audio data, and a represents the noise energy adjustment ratio. If the simulation noisy data x(t) in formula 9 is taken as the loudspeaker input signal, then x(t) in formula 9 also represents the loudspeaker input signal.

[0100] For the loudspeaker input signal x(t), the change after the audio passes through the saturation region of the power amplifier inside the loudspeaker can be simulated by formula (10):

[0101]

[0102] where x(t) represents the loudspeaker input signal. x max represents the maximum value of the input speech signal x(t). represents the loudspeaker power amplifier audio that has changed after passing through the saturation region of the power amplifier inside the loudspeaker.

[0103] In some embodiments, the loudspeaker power amplifier audio can also be generated according to the clean speech audio data, and then x(t) in formula (10) is the clean speech audio data s(t) in formula (9).

[0104] In some embodiments, the loudspeaker power amplifier audio can also be generated according to the noise audio data, and then x(t) in formula (10) is the noise audio data n(t) in formula (9).

[0105] Further, the loudspeaker power amplifier audio is subjected to a first nonlinear conversion to obtain nonlinear loudspeaker power amplifier audio. The first nonlinear conversion can be represented by formula (11):

[0106]

[0107] where x(t) represents the loudspeaker input signal.

[0108] Further, the nonlinear loudspeaker power amplifier audio is processed using a nonlinear action function to generate simulation loudspeaker audio.

[0109] Specifically, the nonlinear characteristics of the loudspeaker can be described in mathematical language by a nonlinear action function sigmoid function, which can be represented by formulas (12)-(13):

[0110]

[0111] where, represents the nonlinear loudspeaker power amplifier audio, represents the loudspeaker power amplifier audio, a is a nonlinear parameter, when a can be 4, when a can be 2.

[0112] represents the simulation loudspeaker audio after the distortion change caused by the nonlinear characteristics of the loudspeaker.

[0113] Thus, the distortion phenomenon of the speech when passing through the saturation area of the power amplifier inside the loudspeaker and the nonlinear change occurring in the transmission process can be simulated by the above formulas (10)-(13). The changes of the speech after the spatial transmission through the loudspeaker are described by mathematical forms, so that the simulation loudspeaker audio can be obtained, which can be stored as independent simulation audio data and can be applied to the scenario of simulating loudspeaker audio.

[0114] When the simulation loudspeaker audio is obtained, the reverberation audio data can be generated according to the simulation loudspeaker audio and the room impulse response.

[0115] wherein the room impulse response (RIR) can be used to realize a specific required RIR signal by using a mirror sound source model and the like.

[0116] Specifically, the simulation loudspeaker audio played by the loudspeaker is convolved with a randomly selected room impulse response signal RIR(t) to generate a signal d(t) with reverberation, and the synthesis can use a known convolution formula, as shown in formula (14):

[0117]

[0118] Thus, the simulation loudspeaker audio obtained by describing the changes of the speech after the spatial transmission through the loudspeaker by mathematical forms, and then convolving the simulation loudspeaker audio with a specific room impulse response, can simulate the synthesized reverberation audio data.

[0119] Optionally, please refer to Figure 2 , the target audio data includes the reverberation audio data, and step 130: generating target audio data for simulating changes of audio after spatial transmission according to the original audio data or the simulation noise data, can include:

[0120] Step 210, generating the simulation echo near-end audio data according to at least one of the simulation noise data, the pure speech audio data and the noise audio data.

[0121] Please refer to Figure 3 , Figure 3To generate the schematic diagram of the echoed audio data, the voice x(n) of the far-end speaker is transmitted through the communication and played out from the near-end speaker A. After the audio is played out by the near-end speaker A, it is transmitted through the near-end environment and collected by the near-end microphone B. Meanwhile, the voice s(n) of the near-end speaker and the noise v(n) possibly existing in the near-end environment are also collected by the near-end microphone B to generate the echoed audio data y(n).

[0122] Specifically, at least one of the simulated noisy data, the pure voice audio data and the noise audio data is used. For example, the simulated noisy data is used. The simulated noisy data can be generated by the above formula (1)-

[0123] (8) to generate the simulated echoed near-end audio data. The method for generating the simulated echoed near-end audio data can be implemented by the above (9)-

[0124] (13). Details are not described herein.

[0125] At step 220, the near-end audio data is convoluted with the room impulse response to generate the simulated echoed near-end reverberation audio.

[0126] Specifically, the near-end audio data played out by the near-end speaker is convoluted with the randomly selected room impulse response signal RIR(n) to generate the simulated echoed near-end reverberation audio d(n). Details are shown in the following formula (15):

[0127]

[0128] At step 230, the echoed audio data is generated according to the near-end reverberation audio and the near-end audio data.

[0129] Specifically, the near-end audio data u(n) includes the voice s(n) of the near-end speaker and the noise v(n) possibly existing in the near-end environment. The near-end voice s(n) and the noise v(n) possibly existing in the near-end environment can be synthesized according to the signal-to-noise ratio (SNR) synthesis method to generate the near-end audio data u(n). Details are shown in the following formula (16):

[0130] u(n) = s(n) + p * v(n) (16);

[0131] wherein the parameter p represents the noise adjustment ratio of the noise audio in the synthesis of the near-end audio data. The calculation method of the parameter p can refer to the above formula (2)-(3). Details are not described herein.

[0132] Further, the echoed audio data e(n) is generated according to the near-end reverberation audio d(n) and the near-end audio data u(n). Details are shown in the following formula (17):

[0133] e(n) = u(n) + q * d(n) (17)

[0134] wherein u(n) refers to formula (16), d(n) refers to formula (15), and parameter q represents an echo audio adjustment ratio when synthesizing the echo audio according to the signal-to-echo ratio (SER), and different echo audio data with different echo levels can be obtained by adjusting q.

[0135] Thus, the echo audio data can be synthesized by the above formulas (1)-(17). Meanwhile, the echo audio data in different application scenarios can also be synthesized by the above formulas (1)-(17), and at least the following scenarios can be included:

[0136] (1) the far end has pure speech and no noise, and the near end has no pure speech;

[0137] (2) the far end has noisy speech, and the near end has no pure speech;

[0138] (3) the far end has no speech and no noise, and the near end has noisy speech;

[0139] (4) the far end has no speech and no noise, and the near end has pure speech;

[0140] (5) the far end has pure speech, and the near end has pure speech;

[0141] (6) the far end has pure speech, and the near end has noisy speech;

[0142] (7) the far end has noisy speech, and the near end has noisy speech.

[0143] wherein for scenario (1), the far end has speech and no noise, and the loudspeaker input pure speech x(n) = s(n). The near end has no pure speech, and thus u(n) = p * v(n) in formula (16).

[0144] For scenario (2), the far end has noisy speech,

[0145] and the loudspeaker input speech x(n) = s(n) + a * n(n). The near end has no pure speech, and thus u(n) = p * v(n) in formula (16).

[0146] For scenario (3), the far end has noisy speech,

[0147] and the loudspeaker input speech x(n) = s(n) + a * n(n). The near end has noisy speech, and thus u(n) = s(n) + p * v(n) in formula (16).

[0148] Similarly, the echo audio data of scenarios (1)-(7) above and the echo audio data of other more application scenarios can be obtained.

[0149] Thus, the echo phenomenon of the speech passing through the near-end loudspeaker and the near-end propagation environment into the near-end microphone can be simulated by the above formulas (1)-(17). Thus, the change of the speech after passing through the echo to generate the spatial transmission is described by the mathematical form, and then various types of simulated echo audio data are synthesized, so that a large amount of data and various types of echo audio data do not need to be artificially prepared and collected. Meanwhile, the diversity of the echo audio data is effectively improved.

[0150] Optionally, in step 230, the step of generating echo audio data according to the near-end reverberation audio and the near-end audio data comprises:

[0151] The near-end reverberation audio of the simulated echo is processed by time delay to obtain the simulated reverberation audio collected by the near-end microphone;

[0152] The simulated reverberation audio collected by the near-end microphone and the near-end audio data are processed according to the signal-to-noise ratio to generate the echo audio data.

[0153] Specifically, the near-end reverberation audio d(n) of the simulated echo is played from the near-end loudspeaker, and needs a certain time delay t delay . Specifically, as shown in formula (18):

[0154]

[0155] wherein, represents the simulated reverberation audio collected by the near-end microphone after the time delay processing. d(n) represents the near-end reverberation audio of the simulated echo, which can be referred to formula (15).

[0156] Further, the simulated reverberation audio collected by the near-end microphone and the near-end audio data are processed according to the signal-to-noise ratio to generate the echo audio data. That is, the near-end audio data u(n) and the reverberation signal According to the randomly selected SER in a certain range, i.e., the parameter q, the final echo audio data is synthesized Specifically, as shown in formula (19).

[0157]

[0158] wherein, the parameter q represents the echo audio scaling coefficient when the echo audio is synthesized according to the SER, t delay represents the time of the reverberation audio signal d(n) passing through the near-end environment, which can be selected as a suitable value in the time range of 0-100 ms.

[0159] In step 140, a speech enhancement operation is performed to obtain enhanced target audio data.

[0160] The speech enhancement includes first-order speech enhancement and second-order speech enhancement, and specifically:

[0161] Before the target audio data is input into the speech model, the first-order speech enhancement operation is performed on the original audio data to obtain first-order speech enhanced original audio data, and the first-order speech enhanced original audio data is used to generate updated target audio data. The first-order speech enhancement at least includes audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement.

[0162] Inside the speech model, random information loss processing is performed on the feature dimension in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data.

[0163] The target audio data generated by the above formulas (1)-(19) can be further speech enhanced to obtain more diversified enhanced target audio data on the basis of the target audio data. The speech enhancement includes at least one first-order speech enhancement and / or at least one high-order speech enhancement.

[0164] As shown in Figure 5 The speech enhancement includes first-order speech enhancement, and the step 140 of performing the speech enhancement operation to obtain the enhanced target audio data includes:

[0165] Before the target audio data is input into the speech model, the first-order speech enhancement operation is performed on the original audio data to obtain first-order speech enhanced original audio data, and the first-order speech enhanced original audio data is used to generate updated target audio data. The first-order speech enhancement at least includes audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement.

[0166] The speech model includes various target audio data generated by the audio data processing method of the present application, and a task model for related speech processing, which can be an AI machine learning model or a non-AI machine learning model, such as a speech processing filter. The speech model can include an AI noise reduction model, an AI echo cancellation model, an AI speech recognition model, a speaker recognition model, etc.

[0167] Before the target audio data is input into the speech model, the first-order speech enhancement operation is performed on the original audio data to obtain first-order speech enhanced original audio data, and the first-order speech enhanced original audio data is used to generate updated target audio data. The first-order speech enhanced original audio data includes first-order speech enhanced pure speech audio data and first-order speech enhanced noise audio data.

[0168] The first-order speech enhancement at least includes audio speed variation, volume adjustment, random displacement, noise enhancement and multiplication enhancement.

[0169] In the formula, the audio speed variation can be achieved by randomly selecting a speed variation coefficient to accelerate or decelerate the original audio data. For example, the original input audio data x(n) is the original audio data, and a speed variation coefficient speed is randomly selected between a maximum speed variation value and a minimum speed variation value. For the acceleration operation with speed>1, the fixed interval point can be taken. For the deceleration operation with speed<1, the first-order linear interpolation can be taken.

[0170] The volume enhancement can be achieved by calculating the volume gain through the exponential distribution. For example, the original input audio data x(n) is the original audio data, and a volume gain range Uniform(min_dBFS, max_dBFS) is set. The volume gain is calculated under the exponential distribution, and the specific formula is shown in formula (20):

[0171]

[0172] β∈Uniform(min_dBFS,max_dBFS)

[0173] The noise enhancement can be achieved by randomly selecting a plurality of noise data noise1(n), noise2(n), … from a noise data set, and then superimposing the selected noise data in the time dimension. 2(n)

[0174] The random displacement enhancement can be achieved by randomly displacing the original audio data. For example, the original input audio data x(n) is the original audio data, and the enhanced audio after random displacement can be represented by formula (21):

[0175] shift aug =x(n-t) (21);

[0176] In the formula, t represents the length of the audio of the random displacement.

[0177] The multiplication enhancement can be used to simulate the speech fluctuation that may exist when a person actually speaks. For example, the original input audio data x(n) is the original audio data, and the original input audio data x(n) is multiplied by a coefficient α, as shown in formula (22):

[0178] aug x(n) =x(n)·α (22);

[0179] In the formula, the coefficient α is subject to a normal distribution, for example, α∈N(0, 1).

[0180] ​Thus, before inputting the target audio data into the speech model, a first-order speech enhancement operation is performed on the original audio data to obtain first-order speech enhanced original audio data. During the first-order speech enhancement operation, the above-mentioned any enhancement method can be enhanced at least once or multiple times, and any combination of multiple first-order speech enhancement methods can also be performed.

[0181] The multi-step first-order audio data enhancement realized through the above-mentioned formulas (20)-(22) performs a nonlinear transformation on the first-order speech enhanced original audio data in the time domain, which can be expressed as formula (23):

[0182] y=F(x(n)) (23);

[0183] Wherein, F(x(n)) represents the first-order speech enhanced original audio data obtained after the combination of any first-order speech enhancement method as mentioned above.

[0184] Thus, the original audio data is subjected to multiple types of first-order speech enhancement operations, which can directly act on the clean speech audio data and noise audio data in the original audio data to batch generate basic and multiple types of audio data.

[0185] Optionally, the method further includes: generating updated simulated noisy data according to the clean speech audio data and noise audio data in the first-order speech enhanced original audio data; and generating updated target audio data according to the first-order speech enhanced original audio data or the updated simulated noisy data.

[0186] Wherein, the updated simulated noisy data is generated according to the clean speech audio data and noise audio data in the first-order speech enhanced original audio data. For example, the multiplicative noise audio data in the first-order speech enhanced noise audio data is converted into updated additive noise audio data through homomorphic filtering processing; and the first-order speech enhanced clean speech audio data and the updated additive noise audio data are synthesized according to the signal-to-noise ratio to obtain the updated simulated noisy data.

[0187] Wherein, the method of generating the updated simulated noisy data can be realized through the above-mentioned formulas (1)-(8), which will not be expanded here.

[0188] Wherein, the updated target audio data is generated according to the first-order speech enhanced original audio data or the updated simulated noisy data.

[0189] For example, the updated target audio data includes updated reverberation audio data. Specifically, according to at least one of the updated simulated noisy audio data, the first-order speech enhanced clean speech audio data, and the first-order speech enhanced noise audio data, an updated simulated speaker audio used to simulate changes of audio after passing through a speaker is generated; and then, according to the updated simulated speaker audio and a room impulse response, the updated reverberation audio data is generated.

[0190] Optionally, generating the updated simulated speaker audio used to simulate changes of audio after passing through a speaker according to at least one of the updated simulated noisy audio data, the first-order speech enhanced clean speech audio data, and the first-order speech enhanced noise audio data includes: processing at least one of the updated simulated noisy audio data, the first-order speech enhanced clean speech audio data, and the first-order speech enhanced noise audio data as a speaker input signal to obtain an updated audio signal maximum value.

[0191] According to the updated audio signal maximum value and the speaker input signal, an updated speaker power amplifier audio used to simulate changes of audio after passing through a power amplifier saturation region inside the speaker is generated.

[0192] Performing a first non-linear conversion on the updated speaker power amplifier audio to obtain an updated non-linear speaker power amplifier audio.

[0193] Processing the updated non-linear speaker power amplifier audio using a non-linear action function to generate the updated simulated speaker audio.

[0194] The method for generating the updated target audio data can be implemented by the above formulas (9)-(19), which will not be described in detail here.

[0195] As shown in FIG. 13, the speech enhancement includes second-order speech enhancement, and the step 140 of performing a speech enhancement operation to obtain enhanced target audio data includes: Figure 5

[0196] In the speech model, the feature dimension of the target audio data and / or the updated target audio data in the time-frequency domain is randomly lost and processed to obtain second-order enhanced target audio data.

[0197] The speech model includes a task model for related speech processing using various audio data generated by the audio data processing method according to the present application, which can be an AI machine learning model or a non-AI machine learning model, such as a speech processing filter. The speech model can include an AI noise reduction model, an AI echo cancellation model, an AI speech recognition model, a speaker recognition model, etc.

[0198] ​The target audio data, or the updated target audio data, or the target audio data and the updated target audio data, is randomly lost in the feature dimension in the time-frequency domain. The following embodiments are described with the target audio data as the input.

[0199] Specifically, the second-order speech enhancement enhances the feature dimension of the target audio data time-frequency graph. In the transmission process of the speech model, the two-dimensional target audio (B, T) is input, B represents the number of audio samples, and T represents the length of the audio data. The two-dimensional target audio (B, T) can be processed by a windowed speech signal, and the two-dimensional target audio (B, T) is converted into three-dimensional (B, T, C) time-frequency domain data.

[0200] Further, the time-frequency unit information of the three-dimensional audio time-frequency domain data is randomly lost in the time domain or the frequency domain. The lost size can be a preset size. The lost part of the three-dimensional audio feature can be filled with 0.

[0201] Please refer to Figure 4 , Figure 4 A feature diagram after a random time-frequency unit loss is shown. The vertical black part represents the randomly lost time domain information, and the horizontal black part represents the randomly lost frequency domain information.

[0202] Optionally, the speech enhancement includes high-order speech enhancement. The step of performing speech enhancement operation to obtain enhanced target audio data includes:

[0203] In the transmission process of the speech model, the second-order enhanced target audio data is randomly lost in the feature dimension in the time-frequency domain to obtain high-order enhanced target audio data.

[0204] Specifically, the target audio data subjected to the second-order speech enhancement in the model can be subjected to multiple random loss information processing to realize the operation of high-order speech enhancement. That is, the high-order speech enhancement operation can include multiple speech enhancement operations, and the speech model can include three-order speech enhancement, four-order speech enhancement, …, N-order speech enhancement.

[0205] It should be noted that the high-order speech enhancement can select the enhancement order according to the actual model structure or business needs. Each order of speech enhancement in the high-order speech enhancement is randomly lost information processing, but the parameters of each order, such as the window function parameters of the windowing, the frequency domain loss, and the time domain loss, can be determined according to the actual model structure or business needs.

[0206] Optionally, the method of randomly losing information in the feature dimension in the time-frequency domain at least once includes:

[0207] windowing and frame shifting processing is performed on the target audio data of the second order enhancement to obtain corresponding three-dimensional audio data;

[0208] randomly losing data of a predetermined range of the time domain and / or the frequency domain of the three-dimensional audio data, so that data of the time domain and / or the frequency domain of the three-dimensional audio data is discontinuous;

[0209] determining target audio data of the high order enhancement according to the three-dimensional audio data after the random loss.

[0210] Specifically, the windowing and frame shifting processing can be performed on the target audio data by using an existing windowing type. The target audio data can include target audio data without speech enhancement, target audio data after the first order speech enhancement, or target audio data after the second order speech enhancement. The following description is given by taking the random loss of information processing on the target audio data as an example.

[0211] For example, two-dimensional target audio data (B, T) is input to a speech model for analysis, where B represents the number of audio samples, and T represents the length of the audio data. The two-dimensional target audio data (B, T) is converted into three-dimensional representation (B, T, C) through windowing and frame shifting processing. For example, the windowing and frame shifting processing is performed on target audio data (B, T) = (1, 16000) with a frame length of 640 and a frame shift of 160. If the first and last frames are not considered, T = 100 frames and C = 640 are obtained, and then the three-dimensional audio is (1, 100, 640).

[0212] Further, the three-dimensional time-frequency domain feature of the three-dimensional audio (B, T, C) is represented as f(x, y, z), and the random loss process in the time dimension is represented by formula (24):

[0213] f(x, y1: y1+Δy, z) = 0 (24);

[0214] where y1 is randomly selected as the loss starting point. Δy is randomly selected within a certain range, and the reference range can be (0-30).

[0215] The random loss process in the frequency dimension is represented by formula (25):

[0216] f(x, y, z1: z1+Δz) = 0 (25);

[0217] where z1 is randomly selected as the loss starting point. Δz is randomly selected within a certain range, and the reference range can be (0-30).

[0218] It should be noted that if the speech model learns and extracts features in the time domain, random loss processing can be directly performed on the three-dimensional (B, T, C) after windowed frame shift processing; if the model extracts features in the time-frequency domain, the three-dimensional (B, T, C) data after windowed frame shift can be Fourier transformed and then subjected to random loss processing.

[0219] Optionally, the time-frequency information of the three-dimensional audio data can be globally randomly lost with a random area size. This process can be expressed as formula (26):

[0220] f(x,y1:y1+Δy,z1:z1+Δz)=0 (26);

[0221] The expressions of parameters y1, z1, Δy, and Δz are the same as those in formulas (24) and (25) and will not be repeated here.

[0222] Optionally, similarly, the two-dimensional audio data can be converted into a four-dimensional feature representation through windowing. The random information loss operation in the four-dimensional speech feature can refer to the above formula (24):

[0223] (26).

[0224] See also Figure 5 , Figure 5 This is a flow chart of an embodiment of speech enhancement operation, in which pure speech audio data ( Figure 5 Clean audio) and noisy audio data ( Figure 5 The clean speech audio data and the noisy audio data after the first-order speech enhancement are then aliased according to the signal-to-noise ratio (SNR), as shown in formulas (2)-(4).

[0225] Furthermore, the simulated noisy data after aliasing, the clean speech audio data after first-order speech enhancement, and the noisy audio data after first-order speech enhancement may be subjected to first-order speech enhancement again before being input into the speech model.

[0226] Once the data enters the speech model, the model performs second-order speech enhancement and multiple higher-order enhancement operations, such as third-order speech enhancement, and N-order speech enhancement, on the aliased simulated noisy data, the clean speech audio data after first-order speech enhancement, and the noisy audio data after first-order speech enhancement. This generates target audio data after second-order speech enhancement, target audio data after third-order speech enhancement, and target audio data after N-order speech enhancement.

[0227] In one example, the audio data processing method of this application was used in an AI noise reduction model, resulting in improvements in the PESQ key indicator at different signal-to-noise ratios. The indicators are detailed in the table below:

[0228]

[0229] The voice quality perceptual evaluation index PESQ (Perceptual Evaluation of Speech Quality) is an objective and full-reference voice quality evaluation method. The algorithm requires a noise-attenuated signal and an original reference signal, can provide a subjective prediction value for objective voice quality evaluation, and the score is between -0.5 and 4.5. The higher the score, the better the voice quality. As can be seen from the above experimental results, the processing effect of the AI noise reduction model after applying the self-audio data processing method is improved at different signal-to-noise ratios. In the table, the model uses signal-to-noise ratios of 0 dB, 5 dB, 10 dB, and 15 dB. The "original AI model" is the PESQ value without using the high-order speech enhancement method. The "N-order speech enhancement + AI model" is the PESQ value using N-order speech enhancement in the AI model.

[0230] In this way, compared with the first-order audio enhancement method directly acting on the original target audio data, the feature dimension of the audio data in the time-frequency domain is randomly lost in the model. On the one hand, the random loss can be controlled by the model parameters, which is closely related to the actual required voice processing business. According to different voice processing businesses, the corresponding random loss effect is performed. On the other hand, the diversity of the model input voice can be effectively increased. If the input model is an audio without losing information, the model will excessively rely on the complete context relationship of the audio. When the information is randomly lost, the model can be forced to pay attention to the relationship between the audio at a slightly distant interval, learn more information from the data, and improve the performance of the model. At the same time, the first-order and high-order enhancement operations, as well as the higher-order data enhancement strategy, further improve the diversity of the data set and the generalization ability of the model, and the original target audio data is expanded more complexly.

[0231] All the above technical methods can be combined to form optional embodiments of the present application, which will not be described one by one here.

[0232] Optionally, the simulation audio data set is constructed according to at least one of the pure voice audio data, the noise audio data, the simulated noise data, and the target audio data.

[0233] Specifically, the simulation audio data set is constructed according to the simulated noise data and the target audio data generated by any of the above embodiments, and the target audio data after voice enhancement. Or, the simulation audio data set can be constructed, which contains the simulated noise data and the target audio data obtained in any of the above embodiments, and the original audio data collected, including pure voice audio data and noise audio data.

[0234] Please refer to Figure 6 , Figure 6 is an example of the composition of the simulation audio data set. In which, the simulation audio data set includes collected clean speech audio data and noise audio data, room impulse response, and special consideration scene data, including whisper audio collected for the whispering scene noise reduction problem, and pure music audio data collected for the music noise reduction problem, and the noise audio data and echo audio data synthesized according to the clean speech audio data and noise audio data, room impulse response, and special consideration scene data can be stored as the simulation audio data set. In some embodiments, the synthesis can be performed in real time in actual business applications.

[0235] When it is necessary to use the audio data in the simulation audio data set for voice processing business, voice processing can be performed based on the data in the simulation audio data set, and high-order voice enhancement operations can be performed in the voice processing process to complete the corresponding voice processing task.

[0236] The embodiment of the application obtains collected original audio data, the original audio data including pure speech audio data and noise audio data; generates simulated noise data according to the pure speech audio data and the noise audio data in the original audio data; generates target audio data used for simulating changes of audio after space transmission according to the original audio data or the simulated noise data, wherein the space transmission includes transmission through a loudspeaker or transmission through a near-end loudspeaker and a microphone; performs speech enhancement operation to obtain enhanced target audio data, wherein the speech enhancement includes first-order speech enhancement and second-order speech enhancement, specifically: performing first-order speech enhancement operation on the original audio data before inputting the target audio data into a speech model to obtain first-order speech enhanced original audio data, and the first-order speech enhanced original audio data is used for generating updated target audio data; the first-order speech enhancement includes audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement; performing random information loss processing on the target audio data and / or the updated target audio data in a feature dimension in a time-frequency domain inside the speech model to obtain second-order enhanced target audio data. A large amount of clean human voice audio and various types of noise audio are used to describe changes in the speech space propagation path through mathematical language, and various simulated target audio data are synthesized. Compared with the existing audio data collection by manual collection, which consumes a lot of manpower and material resources, the application uses easily collected original audio data for audio data processing, simulates the transmission and change of audio through various spaces through mathematical language, automatically generates diversified target audio data in batches, and proposes a more complete simulation audio data synthesis method. In addition, speech enhancement operation is proposed for the generated target audio data, the speech enhancement operation includes first-order speech enhancement on the original audio data, and re-generating updated target audio data based on the first-order speech enhanced data, and performing second-order speech enhancement on the target audio data and / or the updated target audio data, further improving the diversity of the data set.

[0237] In addition, compared with the existing data enhancement technology, the speech enhancement operation of the application includes more data enhancement operations, and the speech enhancement operation includes at least one first-order and high-order operation, which effectively improves the diversity of the data set. At the same time, a high-order speech enhancement method is proposed for speech processing tasks based on an AI model such as noise reduction and echo cancellation tasks, which adds high-order enhancement operation to the input model and the model inside on the basis of first-order ordinary audio data enhancement, expands the diversity of the audio data, and improves the generalization ability of the speech model to a certain extent.

[0238] Further, in the AI voice noise reduction and echo cancellation background, the application can construct a simulation audio data set according to at least one of the generated pure voice audio data, noise audio data, simulated noise data and target audio data, and can perform voice processing based on the data in the simulation audio data set to complete the corresponding voice task. At the same time, the original audio data can be more efficiently utilized, the data acquisition cost is effectively reduced, the data utilization rate is maximized, and the performance of the AI voice model in the downstream task is improved.

[0239] In order to better implement the audio data processing method of the embodiments of the application, the embodiments of the application also provide an audio data processing device. Please refer to Figure 7 , Figure 7 The structure diagram of the audio data processing device provided by the embodiments of the application. Wherein, the audio data processing device 700 can include:

[0240] The acquisition unit 710 is configured to acquire the collected original audio data, and the original audio data includes pure voice audio data and noise audio data;

[0241] The generation unit 720 is configured to generate simulated noise data according to the pure voice audio data and the noise audio data in the original audio data; and

[0242] According to the original audio data or the simulated noise data, target audio data used to simulate the change of audio after spatial transmission is generated, wherein the spatial transmission includes transmission through a loudspeaker or transmission through a near-end loudspeaker and a microphone;

[0243] The enhancement unit 730 is configured to perform a voice enhancement operation to obtain enhanced target audio data, wherein the voice enhancement includes first-order voice enhancement and second-order voice enhancement, and specifically:

[0244] Before the target audio data is input into the voice model, a first-order voice enhancement operation is performed on the original audio data to obtain first-order voice enhanced original audio data, and the first-order voice enhanced original audio data is used to generate updated target audio data; the first-order voice enhancement at least includes audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement;

[0245] Inside the voice model, random information loss processing is performed on the feature dimension in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data.

[0246] Optionally, the generation unit 720 can be configured to convert the multiplicative noise audio data in the noise audio data into additive noise audio data through homomorphic filtering processing; and synthesize the pure voice audio data and the additive noise audio data according to the signal-to-noise ratio to obtain the simulated noise data.

[0247] Optionally, the generating unit 720 can also be configured to generate simulated loudspeaker audio according to at least one of the simulated noisy data, the clean speech audio data, and the noise audio data, the simulated loudspeaker audio being used to simulate changes in audio after passing through a loudspeaker; and generate reverberation audio data according to the simulated loudspeaker audio and a room impulse response.

[0248] Optionally, the generating unit 720 can also be configured to process at least one of the simulated noisy data, the clean speech audio data, and the noise audio data as a loudspeaker input signal to obtain an audio signal maximum value; generate loudspeaker power amplifier audio according to the audio signal maximum value and the loudspeaker input signal, the loudspeaker power amplifier audio being used to simulate changes in audio after passing through a power amplifier saturation region inside a loudspeaker; perform first non-linear conversion on the loudspeaker power amplifier audio to obtain non-linear loudspeaker power amplifier audio; and process the non-linear loudspeaker power amplifier audio using a non-linear function to generate the simulated loudspeaker audio.

[0249] Optionally, the generating unit 720 can also be configured to generate near-end audio data of simulated echo according to at least one of the simulated noisy data, the clean speech audio data, and the noise audio data; perform convolution processing on the near-end audio data and the room impulse response to generate near-end reverberation audio of simulated echo; and generate echo-bearing audio data according to the near-end reverberation audio and the near-end audio data.

[0250] Optionally, the generating unit 720 can also be configured to perform delay processing on the near-end reverberation audio of simulated echo to obtain simulated reverberation audio collected by a near-end microphone; and process the simulated reverberation audio collected by the near-end microphone and the near-end audio data according to a signal-to-noise ratio to generate the echo-bearing audio data.

[0251] Optionally, the enhancing unit 730 can be configured to perform at least one random information loss processing on a feature dimension of the second-order enhanced target audio data in a time-frequency domain during data transmission of the speech model to obtain high-order enhanced target audio data.

[0252] Optionally, the enhancing unit 730 can be configured to perform windowing and frame shifting processing on the second-order enhanced target audio data to obtain corresponding three-dimensional audio data; randomly lose data in a predetermined range of a time domain and / or a frequency domain of the three-dimensional audio data, so that data in the time domain and / or the frequency domain of the three-dimensional audio data is discontinuous; and determine the high-order enhanced target audio data according to the three-dimensional audio data after the random loss.

[0253] Optionally, the audio data processing apparatus 700 further comprises a constructing unit 740, which can be configured to construct a simulated audio data set according to at least one of the clean speech audio data, the noise audio data, the simulated noise-added data and the target audio data; and perform speech processing based on data in the simulated audio data set to complete a corresponding speech task.

[0254] Optionally, the enhancing unit 730 can be configured to generate updated simulated noise-added data according to the clean speech audio data and the noise audio data in the first-order speech enhanced original audio data; and generate updated target audio data according to the first-order speech enhanced original audio data or the updated simulated noise-added data.

[0255] It should be noted that the functions of the modules in the audio data processing apparatus 700 in the embodiments of the present application can correspond to the specific implementation manners of any of the method embodiments described above, which will not be described here.

[0256] The various units in the audio data processing apparatus 700 described above can be realized by software, hardware and combinations thereof in whole or in part. The various units can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the various units.

[0257] The audio data processing apparatus 700 may, for example, be integrated in a terminal or a server having a storage and a processor installed and having computing capability, or the audio data processing apparatus 700 is the terminal or the server. The terminal can be a smart phone, a tablet computer, a notebook computer, a smart television, a smart speaker, a wearable smart device, a personal computer (PC) or the like, and the terminal can further include a client, which can be a video client, a browser client or an instant messaging client, etc. The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms.

[0258] Figure 8 The schematic structural diagram of the audio data processing apparatus 800 provided by the embodiments of the present application is shown in Figure 8As shown, the audio data processing apparatus 800 can include a communication interface 801, a memory 802, a processor 803 and a communication bus 804. The communication interface 801, the memory 802 and the processor 803 can communicate with each other through the communication bus 804. The communication interface 801 is configured to perform data communication between the apparatus 800 and an external device. The memory 802 can be configured to store software programs and modules. The processor 803 can execute the software programs and modules stored in the memory 802, such as the software programs corresponding to the operations in the foregoing method embodiments.

[0259] Optionally, the processor 803 can invoke the software programs and modules stored in the memory 802 to perform the following operations:

[0260] obtain the collected raw audio data, the raw audio data including clean speech audio data and noise audio data; generate simulated noisy data according to the clean speech audio data and the noise audio data in the raw audio data; generate target audio data for simulating changes of audio after spatial transmission, according to the raw audio data or the simulated noisy data, wherein the spatial transmission includes transmission through a loudspeaker or transmission through a near-end loudspeaker and a microphone; perform speech enhancement to obtain enhanced target audio data, wherein the speech enhancement includes first-order speech enhancement and second-order speech enhancement, and specifically: performing first-order speech enhancement on the raw audio data before inputting the target audio data into a speech model to obtain first-order speech enhanced raw audio data, and the first-order speech enhanced raw audio data is used to generate updated target audio data; the first-order speech enhancement at least includes audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement; and performing random information loss processing on the target audio data and / or the updated target audio data in a feature dimension in a time-frequency domain inside the speech model to obtain second-order enhanced target audio data.

[0261] Optionally, the audio data processing apparatus 800 can be integrated in a terminal or a server having a storage and being installed with a processor and having a computing capability, or the audio data processing apparatus 800 is the terminal or the server. The terminal can be a smart phone, a tablet computer, a notebook computer, a smart television, a smart speaker, a wearable smart device, a personal computer or the like. The server can be a physical server, a server cluster or a distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and basic cloud computing services such as big data and artificial intelligence platforms.

[0262] Optionally, the present application further provides a computer device, comprising a memory and a processor, the memory stores a computer program, and the processor executes the computer program to realize the steps in the above-mentioned method embodiments.

[0263] The present application further provides a computer readable storage medium for storing a computer program. The computer readable storage medium can be applied to a computer device, and the computer program makes the computer device execute the corresponding procedures in the audio data processing method in the embodiments of the present application. For brevity, details are not repeated here.

[0264] The present application further provides a computer program product, comprising computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the computer device execute the corresponding procedures in the audio data processing method in the embodiments of the present application. For brevity, details are not repeated here.

[0265] The present application further provides a computer program, comprising computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the computer device execute the corresponding procedures in the audio data processing method in the embodiments of the present application. For brevity, details are not repeated here.

[0266] It should be understood that the processor of the embodiments of the present application can be an integrated circuit chip with a processing capability of signals. In the implementation process, each step of the method embodiments described above can be completed by the integrated logic circuit of hardware in the processor or the instructions in the form of software. The processor described above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor or the like. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware coding processor for execution, or a combination of hardware and software modules in the coding processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the storage, and the processor reads the information in the storage, and combines the hardware to complete the steps of the above method.

[0267] It is to be understood that the memory in the embodiments of the present application can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. Among them, the nonvolatile memory can be a read-only memory (Read-Only Memory, ROM), a programmable read-only memory (Programmable ROM, PROM), an erasable programmable read-only memory (Erasable PROM, EPROM), an electrically erasable programmable read-only memory (Electrically EPROM, EEPROM) or a flash memory. The volatile memory can be a random access memory (Random Access Memory, RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (Static RAM, SRAM), dynamic random access memory (Dynamic RAM, DRAM), synchronous dynamic random access memory (Synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (Double Data Rate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (Enhanced SDRAM, ESDRAM), synchronous link dynamic random access memory (Synchlink DRAM, SLDRAM) and direct memory bus random access memory (Direct Rambus RAM, DR RAM). It should be noted that the memory of the system and method described herein is intended to include, but not limited to, these and any other suitable types of memory.

[0268] It should be understood that the above-mentioned memory is exemplary but not limiting, for example, the memory in the embodiments of the present application can also be static random access memory (static RAM, SRAM), dynamic random access memory (dynamic RAM, DRAM), synchronous dynamic random access memory

[0269] (synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (double data rate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (enhanced SDRAM, ESDRAM), synchronous link dynamic random access memory (synch link DRAM, SLDRAM) and direct memory bus random access memory (Direct Rambus RAM, DR RAM) and the like. That is, the memory in the embodiments of the present application is intended to include, but not limited to, these and any other suitable types of memory.

[0270] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical method. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0271] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.

[0272] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0273] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e. they can be located in one place or distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the method of the present embodiment.

[0274] In addition, each functional unit in the embodiments of the present application can be integrated into a processing unit, or each unit can exist physically, or two or more units can be integrated into one unit.

[0275] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical method of the present application or the part of the technical method that essentially contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes various media that can store program codes, such as U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk.

[0276] The above descriptions are only the specific embodiments of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for processing audio data, characterized in that: The method comprises: Acquire collected original audio data, wherein the original audio data includes clean speech audio data and noisy audio data; Generating simulated noisy data according to the clean speech audio data and the noisy audio data in the original audio data; generating target audio data for simulating changes in audio after spatial transmission based on the original audio data or the simulated noisy data, wherein the spatial transmission includes passing through a speaker, or passing through a near-end speaker and a microphone; Perform a speech enhancement operation to obtain enhanced target audio data, wherein the speech enhancement includes first-order speech enhancement and second-order speech enhancement, specifically: Before inputting the target audio data into the speech model, performing the first-order speech enhancement operation on the original audio data to obtain first-order speech-enhanced original audio data, and using the first-order speech-enhanced original audio data to generate updated target audio data; the first-order speech enhancement includes at least audio speed change, volume adjustment, random shift, noise enhancement, and multiplicative enhancement; Within the speech model, random information loss processing is performed on the feature dimensions in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data.

2. The method according to claim 1, characterized in that The generating of simulated noisy data according to the clean speech audio data and the noisy audio data in the original audio data comprises: Converting the multiplicative noise audio data in the noise audio data into additive noise audio data through homomorphic filtering; The clean speech audio data and the additive noise audio data are synthesized according to the signal-to-noise ratio to obtain the simulated noisy data.

3. The method according to claim 1, characterized in that The target audio data includes reverberation audio data, and generating target audio data for simulating changes in audio after passing through a space based on the original audio data or the simulated noisy data includes: generating, based on at least one of the simulated noisy data, the clean speech audio data, and the noisy audio data, simulated loudspeaker audio for simulating changes in audio after passing through a loudspeaker; The reverberant audio data is generated according to the simulated loudspeaker audio and the room impulse response.

4. The method according to claim 3, characterized in that Generating simulated loudspeaker audio for simulating changes in audio after passing through a loudspeaker based on at least one of the simulated noisy data, the clean speech audio data, and the noisy audio data comprises: Processing at least one of the simulated noisy data, the clean speech audio data, and the noisy audio data as a speaker input signal to obtain a maximum audio signal; generating, according to the maximum value of the audio signal and the speaker input signal, a speaker power amplifier audio signal for simulating changes caused by audio passing through a saturation region of a power amplifier inside the speaker; Performing a first nonlinear conversion on the loudspeaker power amplifier audio to obtain nonlinear loudspeaker power amplifier audio; The nonlinear loudspeaker amplifier audio is processed using a nonlinear action function to generate the simulated loudspeaker audio.

5. The method according to claim 1, wherein The target audio data includes audio data with echo, and generating target audio data for simulating changes in audio after spatial transmission based on the original audio data or the simulated noisy data includes: generating near-end audio data of a simulated echo based on at least one of the simulated noisy data, the clean speech audio data, and the noisy audio data; Convolution processing is performed on the near-end audio data and the room impulse response to generate near-end reverberation audio simulating an echo; The audio data with echo is generated according to the near-end reverberation audio and the near-end audio data.

6. The method according to claim 5, characterized in that The generating the audio data with echo according to the near-end reverberation audio and the near-end audio data includes: Delaying the near-end reverberation audio of the simulated echo to obtain the reverberation audio recorded by the simulated near-end microphone; The reverberation audio recorded by the simulated near-end microphone and the near-end audio data are processed according to the signal-to-noise ratio to generate the audio data with echo.

7. The method according to claim 1, characterized in that The speech enhancement further includes high-order speech enhancement, and the performing of the speech enhancement operation to obtain enhanced target audio data further includes: During the data transmission process of the speech model, the second-order enhanced target audio data is subjected to at least one random loss information processing in the feature dimension in the time-frequency domain to obtain high-order enhanced target audio data.

8. The method according to claim 7, characterized in that The performing at least one random loss information processing on the feature dimension in the time-frequency domain includes: Performing windowing and frame shifting processing on the second-order enhanced target audio data to obtain corresponding three-dimensional audio data; Randomly losing data of a predetermined range in the time domain and / or frequency domain of the three-dimensional audio data, so that the data in the time domain and / or frequency domain of the three-dimensional audio data is discontinuous; The high-order enhanced target audio data is determined according to the randomly lost three-dimensional audio data.

9. The method according to any one of claims 1 to 8, characterized in that The method comprises: constructing a simulated audio data set according to at least one of the clean speech audio data, the noisy audio data, the simulated noisy data, and the target audio data; Speech processing is performed based on the data in the simulated audio data set to complete the corresponding speech task.

10. The method according to any one of claims 1 to 8, characterized in that The method further comprises: Generate updated simulated noisy data according to the clean speech audio data and the noisy audio data in the original audio data of the first-order speech enhancement; Updated target audio data is generated according to the first-order speech-enhanced original audio data or the updated simulated noisy data.

11. An audio data processing device, characterized in that: The device comprises: An acquisition unit, configured to acquire collected original audio data, wherein the original audio data includes clean speech audio data and noise audio data; a generating unit, configured to generate simulated noisy data based on the clean speech audio data and the noisy audio data in the original audio data; and generating target audio data for simulating changes in audio after spatial transmission based on the original audio data or the simulated noisy data, wherein the spatial transmission includes passing through a speaker, or passing through a near-end speaker and a microphone; an enhancement unit, configured to perform a speech enhancement operation to obtain enhanced target audio data, wherein the speech enhancement includes first-order speech enhancement and second-order speech enhancement, and the enhancement unit is configured to: Before inputting the target audio data into the speech model, performing the first-order speech enhancement operation on the original audio data to obtain first-order speech-enhanced original audio data, and using the first-order speech-enhanced original audio data to generate updated target audio data; the first-order speech enhancement includes audio speed change, volume adjustment, random shift, noise enhancement, and multiplicative enhancement; Within the speech model, random information loss processing is performed on the feature dimensions in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the method according to any one of claims 1 to 10.

13. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the processor is configured to execute the steps of the method according to any one of claims 1 to 10 by calling the computer program stored in the memory.

14. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Processing method and apparatus for audio data

    CN106024005A

  • Processing method and device for voice recognition in vehicle and electronic equipment

    CN108022591A