Audio data processing method and device, medium, equipment and program product

By generating and enhancing audio data, the problems of high data acquisition costs and insufficient diversity in AI speech model training are solved, efficient and diversified audio data set generation is achieved, and the performance of AI speech model is improved.

CN119993185AActive Publication Date: 2025-05-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510258693.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2021-12-01
Publication Date
2025-05-13
Estimated Expiration
2041-12-01

AI Technical Summary

Technical Problem

The prior art requires a lot of manpower, financial resources and material resources when building the audio database required for AI voice models, and the lack of sufficient quantity or type of training data can easily lead to poor overfitting and recognition effects.

Method used

Improve data diversity by obtaining easy-to-acquire raw audio data, including pure voice audio data and noisy audio data, generate simulated noise data and target audio data, and perform first-order and second-order voice enhancement operations.

Benefits of technology

It realizes the simulation of audio space transmission changes through mathematical language, and automatically generates diversified target audio data in batches, reducing data acquisition costs, and improving the generalization ability and speech processing effect of AI voice models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993185A_ABST
    Figure CN119993185A_ABST
Patent Text Reader

Abstract

The invention discloses an audio data processing method and device, a medium, equipment and a program product, which are applied to scenes such as artificial intelligence (AI), AI noise reduction in machine learning, AI echo cancellation and the like. The method comprises the following steps: acquiring collected original audio data, wherein the original audio data comprises pure voice audio data and noise audio data; according to pure voice audio data and noise audio data in the original audio data, generating simulation noisy data; according to the original audio data or the simulated noisy data, generating target audio data used for simulating the change of the audio after spatial transmission; performing a voice enhancement operation to obtain enhanced target audio data, the voice enhancement operation including performing first-order voice enhancement on the original audio data, regenerating updated target audio data based on the first-order voice enhanced data, and performing second-order voice enhancement on the target audio data and / or the updated target audio data to obtain enhanced target audio data; and the diversity of the simulation audio data set is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese patent application submitted to the China Patent Office on December 1, 2021, with application number 202111456334.6, and invention name "Audio data processing method and device, medium, equipment and program product". Technical Field

[0002] The present application relates to the technical field of speech signal processing, and in particular to an audio data processing method, device, medium, equipment and program product. Background Art

[0003] With the continuous development of speech signal processing technology and artificial intelligence (AI) technology, more and more tasks are being performed on speech through AI, such as AI speech noise reduction and echo cancellation. During the processing, a large amount of audio data such as various noise audios and audios of various special scenes such as echo audios that can be used for training needs to be collected. This collection often consumes a lot of manpower and financial resources. In AI speech processing, if there is a lack of sufficient quantity or type of training data, it is easy to cause problems such as overfitting and poor recognition effect. Summary of the invention

[0004] The embodiments of the present application provide an audio data processing method, apparatus, medium, device and program product, which improve the diversity of simulation audio data synthesis methods.

[0005] On the one hand, a method for processing audio data is provided, the method comprising:

[0006] Acquire collected original audio data, wherein the original audio data includes pure speech audio data and noise audio data;

[0007] Generate simulated noisy data according to the clean speech audio data and the noisy audio data in the original audio data;

[0008] Generate target audio data for simulating changes in audio after passing through a space according to the original audio data or the simulated noisy data, wherein the passing through the space includes passing through a speaker, or passing through a near-end speaker and a microphone;

[0009] Perform a speech enhancement operation to obtain enhanced target audio data, wherein the speech enhancement includes first-order speech enhancement and second-order speech enhancement, specifically:

[0010] Before inputting the target audio data into the speech model, performing the first-order speech enhancement operation on the original audio data to obtain first-order speech-enhanced original audio data, wherein the first-order speech-enhanced original audio data is used to generate updated target audio data; the first-order speech enhancement includes audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement;

[0011] Within the speech model, random loss information processing is performed on the feature dimension in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data.

[0012] In another aspect, an audio data processing device is provided, the device comprising:

[0013] An acquisition unit, used to acquire collected original audio data, wherein the original audio data includes clean speech audio data and noise audio data;

[0014] A generating unit, configured to generate simulated noisy data according to the clean speech audio data and the noisy audio data in the original audio data; and

[0015] Generate target audio data for simulating changes in audio after passing through a space according to the original audio data or the simulated noisy data, wherein the passing through the space includes passing through a speaker, or passing through a near-end speaker and a microphone;

[0016] An enhancement unit is used to perform a speech enhancement operation to obtain enhanced target audio data, wherein the speech enhancement includes first-order speech enhancement and second-order speech enhancement, and the enhancement unit is used to:

[0017] Before inputting the target audio data into the speech model, performing the first-order speech enhancement operation on the original audio data to obtain first-order speech-enhanced original audio data, wherein the first-order speech-enhanced original audio data is used to generate updated target audio data; the first-order speech enhancement includes audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement;

[0018] Within the speech model, random loss information processing is performed on the feature dimension in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data.

[0019] On the other hand, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the audio data processing method of any of the above embodiments.

[0020] On the other hand, a computer device is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and the processor executes the steps of the audio data processing method of any of the above embodiments by calling the computer program stored in the memory.

[0021] On the other hand, a computer program product is provided, comprising computer instructions, which implement the steps of the audio data processing method in any of the above embodiments when the computer instructions are executed by a processor.

[0022] The embodiment of the present application obtains collected original audio data, where the original audio data includes clean speech audio data and noisy audio data; generates simulated noisy data based on the clean speech audio data and the noisy audio data in the original audio data; generates target audio data for simulating changes in audio after spatial transmission based on the original audio data or the simulated noisy data, where spatial transmission includes passing through a speaker, or passing through a near-end speaker and a microphone; performs a speech enhancement operation to obtain enhanced target audio data, where speech enhancement includes first-order speech enhancement and second-order speech enhancement, specifically: before inputting the target audio data into a speech model, performs a first-order speech enhancement operation on the original audio data to obtain first-order speech-enhanced original audio data, and the first-order speech-enhanced original audio data is used to generate updated target audio data; first-order speech enhancement includes audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement; within the speech model, performs random loss information processing on the feature dimension in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data. A large amount of easily available clean human voice audio and various types of noise audio are used to describe the changes in the voice space propagation path through mathematical language, and various simulated target audio data are synthesized. Compared with the related art that consumes a lot of manpower and material resources by manually collecting audio data, this application uses easy-to-collect original audio data for audio data processing, simulates the transmission changes of audio through various spaces through mathematical language, and automatically generates diversified target audio data in batches, proposing a more complete simulated audio data synthesis method. In addition, a speech enhancement operation is proposed for the generated target audio data, and the speech enhancement operation includes performing first-order speech enhancement on the original audio data, and regenerating updated target audio data based on the first-order speech enhanced data, and performing second-order speech enhancement on the target audio data and / or the updated target audio data, further improving the diversity of the simulated audio data set. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical methods in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0024] Figure 1 A flowchart of an audio data processing method provided in an embodiment of the present application;

[0025] Figure 2 Another schematic diagram of a flow chart of an audio data processing method provided in an embodiment of the present application;

[0026] Figure 3 An example diagram of an audio data processing method provided in an embodiment of the present application;

[0027] Figure 4 Another example diagram of the audio data processing method provided in an embodiment of the present application;

[0028] Figure 5 Another example diagram of the audio data processing method provided in an embodiment of the present application;

[0029] Figure 6 Another example diagram of the audio data processing method provided in an embodiment of the present application;

[0030] Figure 7 A schematic diagram of the structure of an audio data processing device provided in an embodiment of the present application;

[0031] Figure 8 A schematic structural diagram of an audio data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical methods in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0033] The embodiments of the present application provide an audio data processing method, an audio data processing device, a medium and a device. Specifically, the method of the embodiments of the present application can be executed by a computer device, wherein the computer device can be a terminal or a server. The embodiments of the present application can be applied to the research of artificial intelligence (AI) noise reduction and artificial intelligence (AI) echo cancellation technology in artificial intelligence and machine learning. At the same time, it can also be used as an auxiliary method to improve data diversity in the research process of technologies such as speech recognition and speaker recognition.

[0034] First, some nouns or terms that appear in the description of the embodiments of the present application are explained as follows:

[0035] Artificial Intelligence (AI): It is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making. The AI ​​speech model in this application processes the speech audio through the AI ​​machine learning model to obtain the corresponding analysis results, such as AI speech noise reduction, echo cancellation and other operations.

[0036] Signal-to-noise ratio (SNR): It indicates the ratio of the amplitude of useful signal to that of noise signal.

[0037] Perceptual Evaluation of Speech Quality (PESQ): PESQ is an objective, full-reference method for evaluating speech quality. Its algorithm requires a noisy attenuated signal and an original reference signal. It can provide a subjective prediction value for the objective speech quality evaluation. The score ranges from -0.5 to 4.5. The higher the score, the better the speech quality.

[0038] Blockchain system: It can be a distributed system formed by connecting clients and multiple nodes (any form of computing devices connected to the network, such as servers and user terminals) through network communication. The nodes form a peer-to-peer (P2P) network. The P2P protocol is an application layer protocol running on the Transmission Control Protocol (TCP). In a distributed system, any machine such as a server or terminal can join and become a node. The node includes the hardware layer, the middle layer, the operating system layer and the application layer.

[0039] The rise of data-driven AI algorithms and their widespread product applications have brought great convenience to our lives. Data-driven AI algorithms enable models to understand data and learn knowledge from historical data to predict unknown data. Therefore, in order to create an AI model with strong generalization capabilities, the machine needs to be knowledgeable and have accumulated sufficient data.

[0040] Among them, building the audio database required for AI voice model training requires a relatively high cost in practice. For example, when using AI voice models to process audio, such as AI voice recognition training, or in actual applications, a large amount of audio data is required for training. Currently, there is a general basic audio database, but the general data diversity is not enough to meet various business needs. It is often necessary to manually collect a large amount of various noise audio, echo audio, and audio data for various special application scenarios that can be used for training. However, this type of collection requires a lot of manpower, financial resources, and material resources.

[0041] In another category of techniques, an audio data augmentation method for speech recognition is proposed. The augmentation strategies include using functional sparse image warping to apply time warping, randomly selecting audio frequency domain channels for masking, and randomly selecting audio time domain channels for masking. However, in some special application scenarios, more diverse datasets are still needed.

[0042] From the inception to the development of AI voice technology, a complete industrial chain including upstream, midstream and downstream has been formed. The downstream industry applications of intelligent voice technology are diversified, and there is a wide demand for one-stop services. Consumer-level application fields currently include but are not limited to: chat apps, smart hardware, smart homes, car systems, etc. This application proposes an audio data processing method for technical fields such as AI voice model training that require a variety of audio data. It can be applied to the research of artificial intelligence, AI noise reduction in machine learning, and AI echo cancellation technology. At the same time, it can also be used as an auxiliary method to improve data diversity in the process of research on technologies such as speech recognition and speaker recognition. It proposes a relatively complete audio data processing method, which effectively improves the diversity of audio data sets.

[0043] In order to better understand the technical method provided in the embodiment of the present application, the following briefly introduces the application scenarios applicable to the technical method provided in the embodiment of the present application. It should be noted that the application scenarios introduced below are only used to illustrate the embodiment of the present application and are not limited. Take the audio data processing method executed by a computer device as an example, where the computer device can be a terminal device or a server.

[0044] The embodiments of the present application can be implemented in combination with cloud technology or blockchain network technology. As disclosed in the embodiments of the present application, the audio data processing method, wherein these data can be stored on the blockchain, for example: original audio data, pure voice audio data, noisy audio data, simulated noisy data, target audio data, and enhanced target audio data, can all be stored on the blockchain.

[0045] In order to facilitate the storage and query of original audio data, pure voice audio data, noisy audio data, simulated noisy data, target audio data, and enhanced target audio data, optionally, the audio data processing method also includes: sending original audio data, pure voice audio data, noisy audio data, simulated noisy data, target audio data, and enhanced target audio data to the blockchain network, so that the nodes of the blockchain network fill the original audio data, pure voice audio data, noisy audio data, simulated noisy data, target audio data, and enhanced target audio data into the new block, and when the new block is consistent with the consensus, the new block is appended to the tail of the blockchain. The embodiment of the present application can store original audio data, pure voice audio data, noisy audio data, simulated noisy data, target audio data, and enhanced target audio data on the chain, realize the backup of the record, when it is necessary to obtain the target audio data and enhanced target audio data, the corresponding target audio data and enhanced target audio data can be directly and quickly obtained from the blockchain, thereby improving the efficiency of audio data processing.

[0046] It should be noted that the order of description of the following embodiments is not intended to limit the priority order of the embodiments.

[0047] Each embodiment of the present application provides an audio data processing method. The embodiments of the present application are described by taking audio data processing by a computer device as an example.

[0048] See also Figure 1 , Figure 1 A flowchart of an audio data processing method provided in an embodiment of the present application, the method comprising:

[0049] Step 110: Acquire collected original audio data, where the original audio data includes pure speech audio data and noise audio data.

[0050] Specifically, the original audio data at least includes pure speech audio data and noise audio data, and may further include room impulse response.

[0051] Among them, the clean voice audio data includes clean human voice audio data. The data type is not limited, and may include multiple languages, multiple dialects, and human voice humming music. The collection of clean human voice audio data can be carried out in advance in a variety of ways, including obtaining from public audio databases, proprietary audio databases, or manually collecting. For example, a large amount of clean human voice audio can be recorded from real scenes.

[0052] Noise audio data includes various types of noise audio data, such as a large amount of pure subway noise, traffic noise, natural background noise, vehicle noise, indoor and outdoor noise, and other common noise scenarios. The collection of noise audio data can be done in advance in a variety of ways, including public audio databases, proprietary audio databases, or manual collection.

[0053] Step 120: Generate simulated noisy data according to the clean speech audio data and the noisy audio data in the original audio data.

[0054] Specifically, after obtaining the pure voice audio data and the noisy audio data. For example, the clean human voice audio data can be added to the noisy audio data to synthesize a new simulated noisy data, and then a large amount of simulated noisy data can be synthesized by aliasing the pure voice audio data with multiple types of noisy audio data. Among them, the simulated noisy data contains a voice signal and its background noise signal. In natural life, a large amount of audio data includes noisy audio data of voice signals and background noise signals. In practical applications of audio, such as the application of AI echo cancellation models, the input of the far-end microphone is human voice and background noise, and the input of the near-end microphone is also human voice and background noise. The input of the far-end microphone and the near-end microphone can be simulated by synthesizing simulated noisy data.

[0055] The method for synthesizing the aliased pure voice audio data and various types of noisy audio data can be implemented according to the signal-to-noise ratio SNR, and the noisy data is synthesized according to the audio signal-to-noise ratio, which is the ratio of the normal sound signal strength to the noise signal strength. For example:

[0056] The signal-to-noise ratio calculation method can be expressed as formula (1):

[0057]

[0058] Where SNR represents the signal-to-noise ratio in dB, s(t) represents pure speech audio data, ∑ t s 2 (t) represents the speech energy of pure speech audio data, n(t) represents the noisy audio data, ∑ t n 2 (t) represents the noise energy of the noisy audio data.

[0059] When synthesizing pure speech audio data and multiple types of noise audio data, the noise energy can be adjusted to α times the original value, that is, αn(t), and the signal-to-noise ratio is expressed as formula (2):

[0060]

[0061] Wherein, q represents the signal-to-noise ratio (SNR) under the adjusted noise energy ratio.

[0062] Therefore, the calculation formula of the noise energy adjustment ratio α can be expressed as formula (3):

[0063]

[0064] When synthesizing pure speech audio data and multiple types of noisy audio data, by giving a preset signal-to-noise ratio SNR, that is, q in formula (3), optionally, the signal-to-noise ratio q can be randomly selected as an integer in the range of (-5, 20). Then, the noise energy adjustment ratio α can be calculated by formula (3), and then the pure speech audio data s(t) and the noisy audio data n(t) are synthesized according to the signal-to-noise ratio to generate simulated noisy data, which can be expressed as formula (4):

[0065] mix(t)=s(t)+αn(t) (4).

[0066] Optionally, the step of generating simulated noisy data according to the clean speech audio data and the noisy audio data in the original audio data further includes:

[0067] The multiplicative noise audio data in the noise audio data is converted into additive noise audio data through homomorphic filtering;

[0068] The pure speech audio data and the additive noise audio data are synthesized according to the signal-to-noise ratio to obtain simulated noisy data.

[0069] Specifically, the collected noise audio data can be additive noise audio data or multiplicative noise audio data. Among them, additive noise audio data and multiplicative noise audio data are two widely used noise types. Additive noise audio data includes thermal noise, shot noise, etc. The relationship between additive noise audio data and signal is addition. Whether there is a signal or not, additive noise audio data exists. Multiplicative noise audio data is generally caused by an imperfect channel. Their relationship with the signal is multiplication. The signal and multiplicative noise audio data exist at the same time. Additive noise audio data can be used to simulate background noise, and multiplicative noise audio data can be used to simulate the time-varying or nonlinearity of the system.

[0070] It can be understood that the additive noise audio data is the interference of noise on speech as the addition of the two signals in the time domain, or from the energy perspective, it is the superposition relationship between background noise and speech in sound intensity, and the two work together to form a noisy speech signal on the microphone.

[0071] Multiplicative noise audio data refers to the relationship between noise and speech in the time domain and the relationship between noise and speech in the frequency domain. Multiplicative noise audio data can be converted into additive noise audio data by transformation, for example, multiplicative noise audio data or convolution noise audio data can be converted into additive noise audio data by homomorphic filtering.

[0072] The conversion into additive noise audio data by homomorphic filtering may include the following steps:

[0073] First, the multiplicative noise audio data can be expressed in the time domain as:

[0074] x(t)=x1(t)*x2(t) (5);

[0075] Wherein, x(t) represents the multiplicative noise audio data, x1(t) is the speech in the multiplicative noise audio data, and x2(t) is the noise in the multiplicative noise audio data.

[0076] Perform Z transform on formula (5) to convert the convolution signal into a product signal, as shown in formula (6):

[0077] Z[x(t)]=X(z)=X1(z)*X2(z) (6).

[0078] Then, logarithmic operations are performed on both sides of formula (6) to convert the multiplication operation into an addition operation, as shown in formula (7):

[0079]

[0080] Then By performing inverse Z transform, the logarithmic z-domain signal can be transformed into a time-domain signal, as shown in formula (8):

[0081]

[0082] In this way, the multiplicative noise audio data x(n) is converted into additive noise audio data

[0083] Furthermore, the pure speech audio data and the additive noise audio data are synthesized according to the signal-to-noise ratio to obtain simulated noisy data.

[0084] The specific synthesis method can be synthesized by the above formulas (1)-(4), that is, after a preset signal-to-noise ratio SNR is given, the above formulas (1)-(4) are used to synthesize specific simulated noisy data.

[0085] Step 130: Generate target audio data for simulating changes in audio after passing through a space based on the original audio data or the simulated noisy data.

[0086] The transmission through space includes transmission through a loudspeaker, or transmission through a near-end loudspeaker and a microphone.

[0087] Specifically, mathematical language can be used to describe the changes in the transmission of audio data in space. The spatial transmission changes may include changes in transmission through a speaker, or changes through a near-end speaker and a microphone. Accordingly, the target audio data may include reverberation audio data, or audio data with echo, etc.

[0088] Optionally, the target audio data includes reverberation audio data, and step 130 includes:

[0089] Step 1301, generating a simulated loudspeaker audio for simulating changes in audio after passing through a loudspeaker according to at least one of the simulated noisy data, the clean speech audio data and the noisy audio data;

[0090] Step 1302: Generate reverberation audio data based on the simulated speaker audio and the room impulse response.

[0091] Optionally, step 1301 includes:

[0092] Processing at least one of the simulated noisy data, the clean speech audio data and the noisy audio data as a speaker input signal to obtain a maximum value of the audio signal;

[0093] According to the maximum value of the audio signal and the speaker input signal, a speaker power amplifier audio for simulating the change of the audio after passing through the saturation area of ​​the power amplifier inside the speaker is generated;

[0094] Performing a first nonlinear conversion on the loudspeaker power amplifier audio to obtain nonlinear loudspeaker power amplifier audio;

[0095] The nonlinear loudspeaker amplifier audio is processed using a nonlinear action function to generate simulated loudspeaker audio.

[0096] Specifically, at least one of the simulated noisy data, the clean speech audio data and the noisy audio data is processed as the speaker input signal x(t) to obtain the maximum value of the audio signal x max . Optional, x max It can also be set as a proportion of the maximum value of the input signal, such as 80% of the maximum value.

[0097] Furthermore, according to the maximum value of the audio signal and the speaker input signal, a speaker power amplifier audio is generated to simulate the change of the audio after passing through the saturation region of the power amplifier inside the speaker. Specifically, the speaker input signal can be simulated noisy data, pure speech audio data or noisy audio data. For example, according to the simulated noisy data, the simulated noisy data can be synthesized by the above formulas (1)-(8), and the simulated noisy data can be expressed as

[0098] x(t)=s(t)+αn(t) (9);

[0099] Where s(t) represents pure speech audio data, n(t) represents noisy audio data, and α represents the noise energy adjustment ratio. If the simulated noisy data x(t) in Formula 9 is used as the speaker input signal, then x(t) in Formula 9 also represents the speaker input signal.

[0100] For the loudspeaker input signal x(t), the change of the audio after passing through the saturation region of the power amplifier inside the loudspeaker can be simulated by formula (10):

[0101]

[0102] Where x(t) represents the speaker input signal. max Represents the maximum value of the input speech signal x(t). Indicates the speaker amplifier audio that changes after passing through the saturation region of the power amplifier inside the speaker.

[0103] In some embodiments, the speaker amplifier audio may also be generated based on the clean speech audio data, and the x(t) in formula (10) is the clean speech audio data s(t) in formula (9).

[0104] In some embodiments, the speaker amplifier audio may also be generated based on the noise audio data, and the x(t) in formula (10) is the noise audio data n(t) in formula (9).

[0105] Furthermore, a first nonlinear transformation is performed on the speaker power amplifier audio to obtain a nonlinear speaker power amplifier audio. The first nonlinear transformation can be expressed as formula (11):

[0106]

[0107] Where x(t) represents the speaker input signal.

[0108] Furthermore, the nonlinear loudspeaker amplifier audio is processed using a nonlinear action function to generate simulated loudspeaker audio.

[0109] Specifically, the nonlinear characteristics of the loudspeaker can be described in mathematical language through the nonlinear action function sigmoid function, which can be expressed as formulas (12)-(13):

[0110]

[0111] in, represents the nonlinear loudspeaker amplifier audio, represents the speaker amplifier audio, a is a nonlinear parameter, when When a can be 4, When , a can be 2.

[0112] Represents the simulated loudspeaker audio after distortion due to the loudspeaker's nonlinear characteristics.

[0113] In this way, the above formulas (10)-(13) can be used to simulate the distortion phenomenon that occurs when the voice passes through the saturation area of ​​the power amplifier inside the speaker, as well as the nonlinear changes that occur during the transmission process. The changes caused by the voice passing through the spatial transmission of the speaker are described in mathematical form, so that simulated speaker audio can be obtained. The simulated speaker audio can be stored as independent simulated audio data and can be applied to the scene of simulating speaker audio.

[0114] After the simulated loudspeaker audio is obtained, reverberation audio data may be generated according to the simulated loudspeaker audio and the room impulse response.

[0115] Among them, the room impulse response (RIR) can use methods such as the image sound source model to achieve a specific required RIR signal.

[0116] Specifically, the simulated speaker audio played by the speaker Convolve with the randomly selected room impulse response signal RIR(t) to generate a signal d(t) with reverberation. The synthesis can use the well-known convolution formula, as shown in formula (14):

[0117]

[0118] In this way, the simulated speaker audio is obtained by mathematically describing the changes caused by the spatial transmission of speech through the speaker, and then the simulated speaker audio is convolved with a specific room impulse response to simulate the synthetic reverberation audio data.

[0119] Optional, see Figure 2 The target audio data includes audio data with echo. Step 130: generating target audio data for simulating changes in audio after passing through a space according to the original audio data or the simulated noisy data may include:

[0120] Step 210: Generate near-end audio data of simulated echo according to at least one of the simulated noisy data, the clean speech audio data and the noisy audio data.

[0121] See also Figure 3 , Figure 3Schematic diagram of audio data with echo. The voice x(n) of the far-end speaker is transmitted through communication and broadcast from the near-end speaker A. After the near-end speaker A broadcasts the audio, it is transmitted through the near-end environment and recorded by the near-end microphone B. At the same time, the voice s(n) of the near-end speaker and the noise v(n) that may exist in the near-end environment are also recorded by the near-end microphone to generate audio data y(n) with echo.

[0122] Specifically, according to at least one of the simulated noisy data, the clean speech audio data and the noisy audio data. For example, according to the simulated noisy data, the simulated noisy data can be obtained by the above formula (1)-

[0123] (8) is synthesized. The method for generating near-end audio data of simulated echo can be performed by the above (9)-

[0124] (13) is implemented and will not be described in detail here.

[0125] In step 220, convolution processing is performed on the near-end audio data and the room impulse response to generate near-end reverberation audio simulating an echo.

[0126] Specifically, the near-end audio data played by the near-end speaker Convolve with the randomly selected room impulse response signal RIR(n) to generate the near-end reverberation audio d(n) of the simulated echo, as shown in formula (15):

[0127]

[0128] Step 230: Generate audio data with echo according to the near-end reverberation audio and the near-end audio data.

[0129] Specifically, the near-end audio data u(n) includes the voice s(n) of the near-end speaker and the noise v(n) that may exist in the near-end environment. The near-end voice s(n) and the noise v(n) that may exist in the near-end environment may be synthesized according to the signal-to-noise ratio SNR synthesis method to generate the near-end audio data u(n), which can be expressed as formula (16):

[0130] u(n)=s(n)+p*v(n) (16);

[0131] The parameter p represents the noise adjustment ratio of the noise audio when synthesizing the near-end audio data. The calculation method thereof can refer to formulas (2)-(3), which will not be elaborated here.

[0132] Furthermore, according to the near-end reverberation audio d(n) and the near-end audio data u(n), the echo audio data e(n) is generated, and the generation method can be expressed as formula (17):

[0133] e(n)=u(n)+q*d(n) (17);

[0134] Wherein, u(n) refers to formula (16), d(n) refers to formula (15), and parameter q represents the echo audio adjustment ratio when synthesizing the echo audio according to the signal to echo ratio (SER). By adjusting q, echo audio data with different echo degrees can be obtained.

[0135] In this way, the audio data with echo can be synthesized by the above formulas (1)-(17). At the same time, the audio data with echo in different application scenarios can also be synthesized by the above formulas (1)-(17), which can at least include the following scenarios:

[0136] (1) There is pure speech and no noise at the far end, and there is no pure speech at the near end;

[0137] (2) There is noisy speech at the far end, but no pure speech at the near end;

[0138] (3) There is no speech or noise at the far end, but there is noisy speech at the near end;

[0139] (4) There is no speech or noise at the far end, but there is pure speech at the near end;

[0140] (5) There is clean speech at the far end and clean speech at the near end;

[0141] (6) There is pure speech at the far end and noisy speech at the near end;

[0142] (7) There is noisy speech at the far end and noisy speech at the near end.

[0143] For scenario (1), there is a person speaking at the far end and there is no noise, and the loudspeaker inputs pure speech x(n)=s(n). There is no pure speech at the near end, so u(n)=p*v(n) in formula (16).

[0144] For scenario (2), there is noisy speech at the far end.

[0145] Then the speaker input speech x(n) = s(n) + αn(n). If there is no pure speech at the near end, then u(n) = p*v(n) in formula (16).

[0146] For scenario (3), there is noisy speech at the far end.

[0147] Then the speaker input speech x(n) = s(n) + αn(n). If there is noisy speech at the near end, then u(n) = s(n) + p*v(n) in formula (16).

[0148] Similarly, the audio data with echo for the above scenarios (1)-(7) and the audio data with echo for other application scenarios can be obtained.

[0149] In this way, the above formulas (1)-(17) can be used to simulate the echo phenomenon of speech passing through the near-end speaker and the near-end propagation environment and entering the near-end microphone. This makes it possible to describe the changes in speech after the spatial transmission generated by the echo in mathematical form, and then synthesize a variety of simulated audio data with echo, so that there is no need to manually prepare and collect a large amount of data and audio data with echo in a variety of echo scenes. At the same time, the diversity of audio data with echo is effectively improved.

[0150] Optionally, step 230: generating audio data with echo according to the near-end reverberation audio and the near-end audio data, comprises:

[0151] Delay processing is performed on the near-end reverberation audio of the simulated echo to obtain the reverberation audio recorded by the simulated near-end microphone;

[0152] The reverberation audio and near-end audio data recorded by the simulated near-end microphone are processed according to the signal-to-noise ratio to generate audio data with echo.

[0153] Specifically, the near-end reverberation audio d(n) of the simulated echo is played from the near-end speaker, and it takes a certain delay time t after it is transmitted in the near-end environment and recorded by the near-end microphone. delay . Specifically, as shown in formula (18):

[0154]

[0155] in, represents the reverberation audio recorded by the simulated near-end microphone after delay processing. d(n) represents the near-end reverberation audio of the simulated echo. d(n) can refer to formula (15).

[0156] Furthermore, the reverberation audio and near-end audio data recorded by the simulated near-end microphone are processed according to the signal-to-noise ratio to generate audio data with echo. According to the SER parameter q randomly selected within a certain range, the final audio data with echo is synthesized The details are as shown in formula (19).

[0157]

[0158] The parameter q represents the echo audio scaling factor when synthesizing echo audio according to SER, and t delay It indicates the time that the reverberation audio signal d(n) is transmitted in the near-end environment. A suitable value can be selected in the time range of 0 to 100 ms.

[0159] Step 140: Perform a speech enhancement operation to obtain enhanced target audio data.

[0160] The speech enhancement includes first-order speech enhancement and second-order speech enhancement, specifically:

[0161] Before inputting the target audio data into the speech model, performing the first-order speech enhancement operation on the original audio data to obtain first-order speech-enhanced original audio data, wherein the first-order speech-enhanced original audio data is used to generate updated target audio data; the first-order speech enhancement at least includes audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement;

[0162] Within the speech model, random loss information processing is performed on the feature dimension in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data.

[0163] The target audio data generated by the above formulas (1)-(19) can be further subjected to speech enhancement, so as to obtain more diverse enhanced target audio data based on the target audio data. The speech enhancement includes at least one first-order speech enhancement and / or at least one high-order speech enhancement.

[0164] like Figure 5 As shown, speech enhancement includes first-order speech enhancement, step 140: performing speech enhancement operation to obtain enhanced target audio data, including:

[0165] Before the target audio data is input into the speech model, a first-order speech enhancement operation is performed on the original audio data to obtain the first-order speech-enhanced original audio data, and the first-order speech-enhanced original audio data is used to generate updated target audio data; the first-order speech enhancement includes at least audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement.

[0166] Among them, the speech model includes a task model for performing relevant speech processing using various target audio data generated by the audio data processing method of this application, which can be an AI machine learning model or a non-AI machine learning model, such as a speech processing filter, etc. The speech model may include an AI noise reduction model, an AI echo cancellation model, an AI speech recognition model, a speaker recognition model, etc.

[0167] Before the target audio data is input into the speech model, a first-order speech enhancement operation is performed on the original audio data to obtain first-order speech-enhanced original audio data, which is used to generate updated target audio data. The first-order speech-enhanced original audio data includes first-order speech-enhanced pure speech audio data and first-order speech-enhanced noisy audio data.

[0168] First-order speech enhancement includes at least audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement.

[0169] Among them, audio speed change can accelerate or decelerate the original audio data by randomly selecting a speed change coefficient. For example, the original input audio data x(n) is the original audio data, and the speed change coefficient speed is randomly selected between the maximum speed change value and the minimum speed change value. For the acceleration operation of speed>1, it can be achieved by taking points at fixed intervals. For the deceleration operation of speed<1, it can be achieved by first-order linear interpolation.

[0170] Volume enhancement can be performed by exponential distribution calculation to enhance the volume. For example, the original input audio data x(n) is the original audio data, the volume gain range Uniform(min_dBFS, max_dBFS) is set, and the volume gain is calculated under the exponential distribution, as shown in formula (20):

[0171]

[0172] β∈Uniform(min_dBFS,max_dBFS)

[0173] Noise enhancement can be achieved by randomly selecting several noise data segments noise1(n), noise 2(n) ,…, and then superimpose the selected noise data in the time dimension.

[0174] Random displacement enhancement can be achieved by randomly displacing the original audio data. For example, the original input audio data x(n) is the original audio data, and the enhanced audio after random displacement can be expressed as formula (21):

[0175] shift aug = x(nt) (21);

[0176] Where t represents the length of the randomly shifted audio.

[0177] Multiplication enhancement can be used to simulate the ups and downs of a person's voice when actually speaking. For example, the original input audio data x(n) is the original audio data, and the original input audio data x(n) is multiplied by the coefficient α, as shown in formula (22):

[0178] aug x(n) =x(n)·α (22);

[0179] Among them, the coefficient α obeys the normal distribution, for example, α∈N(0,1).

[0180] Thus, before the target audio data is input into the speech model, the first-order speech enhancement operation is performed on the original audio data to obtain the original audio data with first-order speech enhancement. During the first-order speech enhancement operation, any of the above enhancement methods can be enhanced at least once, or multiple times, and any combination of multiple first-order speech enhancement methods can also be performed.

[0181] After the multi-step first-order audio data enhancement implemented by the above formulas (20)-(22), the original audio data of the first-order speech enhancement is nonlinearly transformed in the time domain, which can be expressed as formula (23):

[0182] y=F(x(n)) (23);

[0183] Wherein, F(x(n)) represents the original audio data of first-order speech enhancement obtained by combining any of the above first-order speech enhancement methods.

[0184] In this way, various types of first-order speech enhancement operations are performed on the original audio data, which can directly act on the pure speech audio data and noise audio data in the original audio data, and generate basic and multi-type audio data in batches.

[0185] Optionally, the method also includes: generating updated simulated noisy data based on the clean speech audio data and the noisy audio data in the original audio data of the first-order speech enhancement; generating updated target audio data based on the original audio data of the first-order speech enhancement or the updated simulated noisy data.

[0186] According to the clean speech audio data and the noisy audio data in the original audio data of the first-order speech enhancement, the updated simulated noisy data is generated. For example, the multiplicative noise audio data in the noisy audio data of the first-order speech enhancement is converted into the updated additive noise audio data through homomorphic filtering; the clean speech audio data of the first-order speech enhancement and the updated additive noise audio data are synthesized according to the signal-to-noise ratio to obtain the updated simulated noisy data.

[0187] The method for generating updated simulated noisy data can be implemented by the above formulas (1)-(8), which will not be elaborated here.

[0188] Wherein, updated target audio data is generated according to the original audio data or the updated simulated noisy data with first-order speech enhancement.

[0189] For example, the updated target audio data includes updated reverberation audio data. Specifically, based on at least one of the updated simulated noisy data, the first-order speech enhanced clean speech audio data, and the first-order speech enhanced noise audio data, an updated simulated loudspeaker audio for simulating the change of audio after passing through the loudspeaker is generated; then, based on the updated simulated loudspeaker audio and the room impulse response, the updated reverberation audio data is generated.

[0190] Optionally, based on at least one of the updated simulated noisy data, the first-order speech-enhanced clean speech audio data, and the first-order speech-enhanced noisy audio data, generating updated simulated speaker audio for simulating changes in audio after passing through a speaker, comprising: processing at least one of the updated simulated noisy data, the first-order speech-enhanced clean speech audio data, and the first-order speech-enhanced noisy audio data as a speaker input signal to obtain an updated maximum value of the audio signal;

[0191] Generate updated loudspeaker power amplifier audio for simulating changes in audio after passing through a saturation region of a power amplifier inside the loudspeaker according to the updated maximum value of the audio signal and the loudspeaker input signal;

[0192] Performing a first nonlinear transformation on the updated loudspeaker power amplifier audio to obtain updated nonlinear loudspeaker power amplifier audio;

[0193] The updated nonlinear loudspeaker power amplifier audio is processed using a nonlinear action function to generate updated simulated loudspeaker audio.

[0194] Among them, the method for generating updated target audio data can be implemented through the above formulas (9)-(19), which will not be elaborated here.

[0195] like Figure 5 As shown, speech enhancement includes second-order speech enhancement, step 140: performing speech enhancement operation to obtain enhanced target audio data, including:

[0196] Within the speech model, random loss information processing is performed on the feature dimension in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data.

[0197] Among them, the speech model includes a task model for performing relevant speech processing using various audio data generated by the audio data processing method of this application, which can be an AI machine learning model or a non-AI machine learning model, such as a speech processing filter, etc. The speech model may include an AI noise reduction model, an AI echo cancellation model, an AI speech recognition model, a speaker recognition model, etc.

[0198] The target audio data, or the updated target audio data, or the target audio data and the updated target audio data, are subjected to random loss information processing in the feature dimension in the time-frequency domain. The following embodiments are described with the target audio data as input.

[0199] Specifically, the second-order speech enhancement enhances the feature dimensions such as the time-frequency diagram of the target audio data. During the transmission process of the speech model, the two-dimensional target audio (B, T) is input, B represents the number of audio samples, and T represents the length of the audio data. The two-dimensional target audio (B, T) can be processed by windowed speech signals to convert the two-dimensional target audio (B, T) into a three-dimensional (B, T, C) time-frequency domain data representation.

[0200] Furthermore, the time-frequency unit information of the three-dimensional audio time-frequency domain data is randomly lost, and the information is randomly lost in the time domain or the frequency domain, and the loss size can be a preset size. For the part of the random information loss of the three-dimensional audio feature, the lost part information can be filled with 0.

[0201] See also Figure 4 , Figure 4 The figure shows a characteristic diagram after a random time-frequency unit is lost, where the vertical black part represents the randomly lost time domain information and the horizontal black part represents the randomly lost frequency domain information.

[0202] Optionally, the speech enhancement includes high-order speech enhancement, and the step of performing the speech enhancement operation to obtain enhanced target audio data includes:

[0203] During the transmission of the speech model, the second-order enhanced target audio data is subjected to at least one random loss information processing in the feature dimension in the time-frequency domain to obtain high-order enhanced target audio data.

[0204] Specifically, for the target audio data that performs second-order speech enhancement in the model, multiple random loss information processing can be performed to achieve a high-order speech enhancement operation. That is, the high-order speech enhancement operation may include multiple speech enhancement operations, and the speech model may include third-order speech enhancement, fourth-order speech enhancement, and N-order speech enhancement.

[0205] It should be noted that the order of enhancement can be selected according to the actual model structure or business needs. Each order of speech enhancement in high-order speech enhancement is a random loss information process, but the parameters of each order, such as the window function parameters of windowing, frequency domain loss, and time domain loss, can be determined according to the actual model structure or business needs.

[0206] Optionally, the method of performing at least one random loss information processing on the feature dimension in the time-frequency domain includes:

[0207] Performing windowing and frame shifting processing on the second-order enhanced target audio data to obtain corresponding three-dimensional audio data;

[0208] Randomly losing data of a predetermined range in the time domain and / or frequency domain of the three-dimensional audio data, so that the data in the time domain and / or frequency domain of the three-dimensional audio data is discontinuous;

[0209] The target audio data for high-order enhancement is determined according to the three-dimensional audio data after random loss.

[0210] Specifically, the target audio data can be subjected to windowing frame shift processing by existing windowing types. Wherein, the target audio data may include the target audio data without speech enhancement, or the target audio data after first-order speech enhancement, or the target audio data after second-order speech enhancement. The following is described by taking the target audio data being subjected to random loss information processing as an example.

[0211] For example, input two-dimensional target audio data (B, T) to the speech model for analysis, B represents the number of audio samples, and T represents the length of audio data. The two-dimensional target audio data (B, T) is converted into a three-dimensional representation (B, T, C) by windowing and frame shifting. For example, the target audio data (B, T) = (1, 16000) is windowed and frame shifted, with a frame length of 640 and a frame shift of 160. If the beginning and end of the frame are not considered, a frame length of T = 100 frames and C = 640 can be obtained, and the three-dimensional audio is (1, 100, 640).

[0212] Furthermore, the three-dimensional time-frequency domain features of the three-dimensional audio (B, T, C) are expressed as f(x, y, z), and the random loss process in the time dimension is expressed as formula (24):

[0213] f(x,y1:y1+Δy,z)=0 (24);

[0214] The loss starting point y1 is randomly selected, and the loss duration Δy is randomly selected within a certain range, and the reference range may be (0 to 30).

[0215] The random loss process in the frequency dimension is expressed as formula (25):

[0216] f(x,y,z1:z1+Δz)=0 (25);

[0217] The loss start z1 is randomly selected, and the loss duration Δz is randomly selected within a certain range, and the reference range may be (0-30).

[0218] It should be noted that if the speech model learns to extract features in the time domain, the three-dimensional (B, T, C) after the windowed frame shift processing can be directly subjected to random loss processing; if the model extracts features in the time-frequency domain, the three-dimensional (B, T, C) data after the windowed frame shift can be subjected to Fourier transform and then subjected to random loss processing.

[0219] Optionally, the time-frequency information of the three-dimensional audio data may be globally randomly lost with a random area size. This process can be expressed as formula (26):

[0220] f(x,y1:y1+Δy,z1:z1+Δz)=0 (26);

[0221] The expressions of parameters y1, z1, Δy, and Δz are the same as those in formulas (24) and (25) and will not be repeated here.

[0222] Optionally, similarly, the two-dimensional audio data can be converted into a four-dimensional feature representation through windowing. The random information loss operation in the four-dimensional speech feature can refer to the above formula (24):

[0223] (26).

[0224] See also Figure 5 , Figure 5 This is a flow chart of an embodiment of a speech enhancement operation, in which pure speech audio data ( Figure 5 Clean audio) and noisy audio data ( Figure 5 The noise audio shown in the figure can perform first-order speech enhancement. Then, the clean speech audio data and the noise audio data after the first-order speech enhancement are aliased according to the signal-to-noise ratio SNR, which can be referred to formulas (2)-(4).

[0225] Furthermore, for the simulated noisy data after aliasing, the clean speech audio data after first-order speech enhancement and the noisy audio data after first-order speech enhancement, first-order speech enhancement may be performed again before being input into the speech model.

[0226] After entering the speech model, the aliased simulated noisy data, the clean speech audio data after first-order speech enhancement, and the noisy audio data after first-order speech enhancement are subjected to second-order speech enhancement and multiple high-order enhancement operations, such as third-order speech enhancement...N-order speech enhancement, to generate target audio data after second-order speech enhancement, target audio data after third-order speech enhancement...and target audio data after N-order speech enhancement.

[0227] In one example, the audio data processing method of the present application is used in an AI noise reduction model, so that the PESQ key indicators of the AI ​​noise reduction model are improved at different signal-to-noise ratios. The indicators are detailed in the following table:

[0228]

[0229] Among them, the perceptual evaluation index of speech quality, PESQ (Perceptual Evaluation of SpeechQuality, PESQ), is an objective, full-reference speech quality evaluation method. Its algorithm requires a noisy attenuated signal and an original reference signal, which can provide a subjective prediction value for the objective speech quality evaluation. The score is between -0.5 and 4.5. The higher the score, the better the speech quality. As can be seen from the above experimental results, the processing effect of the AI ​​denoising model after applying its own audio data processing method is improved at different signal-to-noise ratios. In the table, the models use signal-to-noise ratios of 0dB, 5dB, 10dB and 15dB respectively. The "original AI model" is the PESQ value without using the high-order speech enhancement method, and the "N-order speech enhancement + AI model" is the PESQ value using N-order speech enhancement in the AI ​​model.

[0230] In this way, compared with the first-order audio enhancement method that directly acts on the original target audio data, the feature dimensions of the audio data in the time-frequency domain are randomly lost within the model. On the one hand, the random loss can be controlled by the model parameters, which is closely related to the actual required voice processing business, and the corresponding random loss effect is performed according to different voice processing businesses. On the other hand, it can effectively increase the diversity of the model input speech, avoiding the problem that if the input model is complete audio without information loss, the model will be overly dependent on the complete contextual relationship of the audio. When information is randomly lost, the model can be forced to pay attention to the relationship between audios that are slightly farther apart, learn more information from the data, and improve the performance of the model. At the same time, the simultaneous first-order and higher-order enhancement operations and higher-order data enhancement strategies further improve the diversity of the data set, improve the generalization ability of the model, and perform more complex expansions on the original target audio data.

[0231] All of the above technical methods can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0232] Optionally, a simulated audio data set is constructed based on at least one of clean speech audio data, noisy audio data, simulated noisy data and target audio data.

[0233] Specifically, a simulated audio data set is constructed based on the simulated noisy data and target audio data generated in any of the above embodiments, and the target audio data after speech enhancement. In other words, a simulated audio data set can be constructed, which includes the simulated noisy data and target audio data obtained in any of the above embodiments, and the collected original audio data, including pure speech audio data and noise audio data.

[0234] See also Figure 6 , Figure 6 This is an example of the composition of a simulated audio data set. The simulated audio data set includes collected pure speech audio data and noise audio data, room impulse response, and data of special consideration scenarios, including whispering audio collected for the mute problem of whispering scenarios, and pure music audio data collected for the mute problem of music. At the same time, noisy audio data and echo audio data synthesized based on pure speech audio data and noise audio data, room impulse response, and data of special consideration scenarios can be stored as simulated audio data sets. In some embodiments, the synthesis can be performed in real time in actual business applications.

[0235] When it is necessary to use the audio data in the simulated audio data set for speech processing services, speech processing can be performed based on the data in the simulated audio data set, and high-order speech enhancement operations can be performed during the speech processing process to complete the corresponding speech processing task.

[0236] The embodiment of the present application obtains collected original audio data, which includes pure speech audio data and noisy audio data; generates simulated noisy data based on the pure speech audio data and the noisy audio data in the original audio data; generates target audio data for simulating changes in audio after spatial transmission based on the original audio data or the simulated noisy data, wherein spatial transmission includes passing through a speaker, or passing through a near-end speaker and a microphone; performs a speech enhancement operation to obtain enhanced target audio data, wherein speech enhancement includes first-order speech enhancement and second-order speech enhancement, specifically: before inputting the target audio data into a speech model, performs a first-order speech enhancement operation on the original audio data to obtain first-order speech-enhanced original audio data, and the first-order speech-enhanced original audio data is used to generate updated target audio data; first-order speech enhancement includes audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement; within the speech model, performs random loss information processing on the feature dimension in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data. A large amount of easily available clean human voice audio and various types of noise audio are used to describe the changes in the voice space propagation path through mathematical language, and various simulated target audio data are synthesized. Compared with the existing manual collection of audio data, which consumes a lot of manpower and material resources, this application uses easy-to-collect original audio data for audio data processing, simulates the transmission changes of audio through various spaces through mathematical language, and automatically generates diversified target audio data in batches, proposing a more complete simulated audio data synthesis method. In addition, a speech enhancement operation is proposed for the generated target audio data, which includes performing first-order speech enhancement on the original audio data, and regenerating updated target audio data based on the first-order speech enhanced data, and performing second-order speech enhancement on the target audio data and / or the updated target audio data, further improving the diversity of the data set.

[0237] In addition, compared to the existing data enhancement technology, which includes more data enhancement operations, the speech enhancement operation of this application effectively improves the diversity of the data set through multi-order speech enhancement operations that include at least one first-order and high-order operation. At the same time, for speech processing tasks based on AI models, such as noise reduction and echo cancellation tasks, a high-order speech enhancement method is proposed. On the basis of first-order ordinary audio data enhancement, high-order enhancement operations are added to the input model and the model, which expands the diversity of audio data and improves the generalization ability of the speech model to a certain extent.

[0238] Furthermore, in the context of AI voice noise reduction and echo cancellation, this application can construct a simulated audio data set based on at least one of the generated pure voice audio data, noisy audio data, simulated noisy data, and target audio data, and can perform voice processing based on the data in the simulated audio data set to complete the corresponding voice task. At the same time, it can more efficiently utilize the original audio data, effectively reduce the data collection cost, maximize data utilization, and improve the performance of the AI ​​voice model in downstream tasks.

[0239] In order to better implement the audio data processing method of the present application embodiment, the present application embodiment also provides an audio data processing device. Figure 7 , Figure 7 A schematic diagram of the structure of an audio data processing device provided in an embodiment of the present application. The audio data processing device 700 may include:

[0240] An acquisition unit 710 is used to acquire collected original audio data, where the original audio data includes pure speech audio data and noise audio data;

[0241] A generating unit 720, configured to generate simulated noisy data according to the clean speech audio data and the noisy audio data in the original audio data; and

[0242] Generate target audio data for simulating changes in audio after passing through a space according to the original audio data or the simulated noisy data, wherein passing through the space includes passing through a speaker, or passing through a near-end speaker and a microphone;

[0243] The enhancement unit 730 is used to perform a speech enhancement operation to obtain enhanced target audio data, wherein the speech enhancement includes first-order speech enhancement and second-order speech enhancement, specifically:

[0244] Before inputting the target audio data into the speech model, a first-order speech enhancement operation is performed on the original audio data to obtain the first-order speech-enhanced original audio data, and the first-order speech-enhanced original audio data is used to generate updated target audio data; the first-order speech enhancement includes at least audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement;

[0245] Within the speech model, random loss information processing is performed on the feature dimension in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data.

[0246] Optionally, the generating unit 720 can be used to convert the multiplicative noise audio data in the noise audio data into additive noise audio data through homomorphic filtering; synthesize the pure speech audio data and the additive noise audio data according to the signal-to-noise ratio to obtain simulated noisy data.

[0247] Optionally, the generation unit 720 can also be used to generate simulated speaker audio for simulating changes in audio after passing through a speaker based on at least one of the simulated noisy data, clean speech audio data and noisy audio data; and generate reverberation audio data based on the simulated speaker audio and the room impulse response.

[0248] Optionally, the generation unit 720 can also be used to process at least one of the simulated noisy data, the clean speech audio data and the noisy audio data as a speaker input signal to obtain a maximum value of the audio signal; generate a speaker power amplifier audio for simulating the change of the audio after passing through the saturation region of the power amplifier inside the speaker according to the maximum value of the audio signal and the speaker input signal; perform a first nonlinear conversion on the speaker power amplifier audio to obtain nonlinear speaker power amplifier audio; and process the nonlinear speaker power amplifier audio using a nonlinear action function to generate simulated speaker audio.

[0249] Optionally, the generation unit 720 can also be used to generate near-end audio data with simulated echo based on at least one of the simulated noisy data, clean speech audio data and noisy audio data; convolve the near-end audio data and the room impulse response to generate near-end reverberation audio with simulated echo; and generate audio data with echo based on the near-end reverberation audio and the near-end audio data.

[0250] Optionally, the generation unit 720 can also be used to delay the near-end reverberation audio of the simulated echo to obtain the reverberation audio recorded by the simulated near-end microphone; and process the reverberation audio and near-end audio data recorded by the simulated near-end microphone according to the signal-to-noise ratio to generate audio data with echo.

[0251] Optionally, the enhancement unit 730 can be used to perform at least one random loss information processing on the feature dimension of the second-order enhanced target audio data in the time-frequency domain during the data transmission process of the speech model to obtain high-order enhanced target audio data.

[0252] Optionally, the enhancement unit 730 can be used to perform windowing and frame shifting processing on the second-order enhanced target audio data to obtain corresponding three-dimensional audio data; randomly lose data in a predetermined range of the time domain and / or frequency domain of the three-dimensional audio data to make the time domain and / or frequency domain data of the three-dimensional audio data discontinuous; and determine the high-order enhanced target audio data based on the three-dimensional audio data after the random loss.

[0253] Optionally, the audio data processing device 700 also includes a construction unit 740, which can be used to construct a simulated audio data set based on at least one of pure speech audio data, noisy audio data, simulated noisy data and target audio data; and perform speech processing based on the data in the simulated audio data set to complete a corresponding speech task.

[0254] Optionally, the enhancement unit 730 can be used to generate updated simulated noisy data based on the clean speech audio data and the noise audio data in the original audio data of the first-order speech enhancement; and generate updated target audio data based on the original audio data of the first-order speech enhancement or the updated simulated noisy data.

[0255] It should be noted that the functions of each module in the audio data processing device 700 in the embodiment of the present application can correspond to the specific implementation methods of any embodiment in the above-mentioned method embodiments, and will not be repeated here.

[0256] Each unit in the audio data processing device 700 may be implemented in whole or in part by software, hardware, or a combination thereof. Each unit may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in a computer device in the form of software, so that the processor may call and execute operations corresponding to each unit.

[0257] The audio data processing device 700 can be integrated in a terminal or server with a storage device and a processor installed therein and having computing power, or the audio data processing device 700 is the terminal or server. The terminal can be a smart phone, a tablet computer, a laptop computer, a smart TV, a smart speaker, a wearable smart device, a personal computer (PC), etc. The terminal can also include a client, which can be a video client, a browser client, or an instant messaging client, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0258] Figure 8 A schematic structural diagram of an audio data processing device 800 provided in an embodiment of the present application is shown in FIG. Figure 8As shown, the audio data processing device 800 may include: a communication interface 801, a memory 802, a processor 803 and a communication bus 804. The communication interface 801, the memory 802, and the processor 803 communicate with each other through the communication bus 804. The communication interface 801 is used for the device 800 to communicate data with an external device. The memory 802 can be used to store software programs and modules, and the processor 803 runs the software programs and modules stored in the memory 802, such as the software programs of the corresponding operations in the aforementioned method embodiments.

[0259] Optionally, the processor 803 may call the software program and module stored in the memory 802 to perform the following operations:

[0260] Acquire the collected original audio data, the original audio data includes pure speech audio data and noisy audio data; generate simulated noisy data according to the pure speech audio data and the noisy audio data in the original audio data; generate target audio data for simulating the change of audio after spatial transmission according to the original audio data or the simulated noisy data, wherein the spatial transmission includes passing through a speaker, or passing through a near-end speaker and a microphone; perform a speech enhancement operation to obtain enhanced target audio data, wherein speech enhancement includes first-order speech enhancement and second-order speech enhancement, specifically: before inputting the target audio data into a speech model, perform a first-order speech enhancement operation on the original audio data to obtain first-order speech-enhanced original audio data, and the first-order speech-enhanced original audio data is used to generate updated target audio data; the first-order speech enhancement includes at least audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement; within the speech model, perform random loss information processing on the feature dimension in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data.

[0261] Optionally, the audio data processing device 800 can be integrated in a terminal or server having a storage device and a processor installed therein and having computing power, or the audio data processing device 800 is the terminal or server. The terminal can be a smart phone, a tablet computer, a laptop computer, a smart TV, a smart speaker, a wearable smart device, a personal computer, or other devices. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0262] Optionally, the present application further provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0263] The present application also provides a computer-readable storage medium for storing a computer program. The computer-readable storage medium can be applied to a computer device, and the computer program enables the computer device to execute the corresponding process in the audio data processing method in the embodiment of the present application, which will not be described in detail for the sake of brevity.

[0264] The present application also provides a computer program product, which includes computer instructions, which are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the corresponding process in the audio data processing method in the embodiment of the present application, which will not be described in detail here for the sake of brevity.

[0265] The present application also provides a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the corresponding process in the audio data processing method in the embodiment of the present application, which will not be described here for the sake of brevity.

[0266] It should be understood that the processor of the embodiment of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiment can be completed by the hardware integrated logic circuit or software instructions in the processor. The above processor can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or the hardware and software modules in the decoding processor are combined and performed. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, and other mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0267] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0268] It should be understood that the above-mentioned memory is exemplary but not restrictive. For example, the memory in the embodiment of the present application may also be a static random access memory (static RAM, SRAM), a dynamic random access memory (dynamic RAM, DRAM), a synchronous dynamic random access memory, or a synchronous dynamic random access memory.

[0269] (synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (double data rate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (enhancedSDRAM, ESDRAM), synchronous link dynamic random access memory (synch link DRAM, SLDRAM) and direct memory bus random access memory (Direct Rambus RAM, DR RAM), etc. That is to say, the memory in the embodiments of the present application is intended to include but not limited to these and any other suitable types of memory.

[0270] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical method. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0271] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0272] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0273] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the method of this embodiment.

[0274] In addition, each functional unit in the embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0275] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical method of the present application, or the part that contributes to the prior art or the part of the technical method, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server) to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks or optical disks.

[0276] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. An audio data processing method, characterized in that: The method comprises: Acquire collected original audio data, wherein the original audio data includes pure speech audio data and noise audio data; Generate simulated noisy data according to the clean speech audio data and the noisy audio data in the original audio data; Generate target audio data for simulating changes in audio after passing through a space according to the original audio data or the simulated noisy data, wherein the passing through the space includes passing through a speaker, or passing through a near-end speaker and a microphone; Perform a speech enhancement operation to obtain enhanced target audio data, wherein the speech enhancement includes first-order speech enhancement and second-order speech enhancement, specifically: Before inputting the target audio data into the speech model, performing the first-order speech enhancement operation on the original audio data to obtain first-order speech-enhanced original audio data, wherein the first-order speech-enhanced original audio data is used to generate updated target audio data; the first-order speech enhancement at least includes audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement; Within the speech model, random loss information processing is performed on the feature dimension in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data.

2. The method according to claim 1, characterized in that The generating of simulated noisy data according to the clean speech audio data and the noisy audio data in the original audio data comprises: Converting the multiplicative noise audio data in the noise audio data into additive noise audio data through homomorphic filtering; The clean speech audio data and the additive noise audio data are synthesized according to the signal-to-noise ratio to obtain the simulated noisy data.

3. The method according to claim 1, characterized in that The target audio data includes reverberation audio data, and generating target audio data for simulating changes in audio after passing through a space according to the original audio data or the simulated noisy data includes: Generate simulated loudspeaker audio for simulating changes in audio after passing through a loudspeaker according to at least one of the simulated noisy data, the clean speech audio data and the noisy audio data; The reverberant audio data is generated according to the simulated speaker audio and the room impulse response.

4. The method according to claim 3, characterized in that The generating, based on at least one of the simulated noisy data, the clean speech audio data and the noisy audio data, a simulated loudspeaker audio for simulating changes in the audio after passing through the loudspeaker comprises: Processing at least one of the simulated noisy data, the clean speech audio data and the noisy audio data as a speaker input signal to obtain a maximum value of the audio signal; Generate, according to the maximum value of the audio signal and the speaker input signal, a speaker power amplifier audio for simulating the change of the audio after passing through the saturation region of the power amplifier inside the speaker; Performing a first nonlinear conversion on the loudspeaker power amplifier audio to obtain nonlinear loudspeaker power amplifier audio; The nonlinear loudspeaker power amplifier audio is processed using a nonlinear action function to generate the simulated loudspeaker audio.

5. The method according to claim 1, characterized in that The target audio data includes audio data with echo, and the step of generating target audio data for simulating changes in audio after passing through a space according to the original audio data or the simulated noisy data includes: Generate near-end audio data of simulated echo according to at least one of the simulated noisy data, the clean speech audio data and the noise audio data; Convolution processing is performed on the near-end audio data and the room impulse response to generate near-end reverberation audio simulating an echo; The audio data with echo is generated according to the near-end reverberation audio and the near-end audio data.

6. The method according to claim 5, characterized in that The generating the audio data with echo according to the near-end reverberation audio and the near-end audio data comprises: Delay processing is performed on the near-end reverberation audio of the simulated echo to obtain the reverberation audio recorded by the simulated near-end microphone; The reverberation audio recorded by the simulated near-end microphone and the near-end audio data are processed according to the signal-to-noise ratio to generate the audio data with echo.

7. The method according to claim 1, characterized in that The speech enhancement further includes high-order speech enhancement, and the performing of the speech enhancement operation to obtain enhanced target audio data further includes: During the data transmission process of the speech model, the second-order enhanced target audio data is subjected to at least one random loss information processing in the feature dimension in the time-frequency domain to obtain high-order enhanced target audio data.

8. The method according to claim 7, characterized in that The performing at least one random loss information processing on the feature dimension in the time-frequency domain comprises: Performing windowing and frame shifting processing on the second-order enhanced target audio data to obtain corresponding three-dimensional audio data; Randomly losing data of a predetermined range in the time domain and / or frequency domain of the three-dimensional audio data, so that the data in the time domain and / or frequency domain of the three-dimensional audio data is discontinuous; The high-order enhanced target audio data is determined according to the three-dimensional audio data after the random loss.

9. The method according to any one of claims 1 to 8, characterized in that: The method comprises: Constructing a simulated audio data set according to at least one of the clean speech audio data, the noisy audio data, the simulated noisy data and the target audio data; Speech processing is performed based on the data in the simulated audio data set to complete a corresponding speech task.

10. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: Generate updated simulated noisy data according to the clean speech audio data and the noisy audio data in the original audio data of the first-order speech enhancement; Updated target audio data is generated according to the first-order speech enhanced original audio data or the updated simulated noisy data.

11. An audio data processing device, characterized in that: The device comprises: An acquisition unit, used to acquire collected original audio data, wherein the original audio data includes clean speech audio data and noise audio data; A generating unit, configured to generate simulated noisy data according to the clean speech audio data and the noisy audio data in the original audio data; and Generate target audio data for simulating changes in audio after passing through a space according to the original audio data or the simulated noisy data, wherein the passing through the space includes passing through a speaker, or passing through a near-end speaker and a microphone; An enhancement unit is used to perform a speech enhancement operation to obtain enhanced target audio data, wherein the speech enhancement includes first-order speech enhancement and second-order speech enhancement, and the enhancement unit is used to: Before inputting the target audio data into the speech model, performing the first-order speech enhancement operation on the original audio data to obtain first-order speech-enhanced original audio data, wherein the first-order speech-enhanced original audio data is used to generate updated target audio data; the first-order speech enhancement includes audio speed change, volume adjustment, random displacement, noise enhancement and multiplication enhancement; Within the speech model, random loss information processing is performed on the feature dimension in the time-frequency domain of the target audio data and / or the updated target audio data to obtain second-order enhanced target audio data.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the method according to any one of claims 1 to 10.

13. A computer device, characterized in that: The computer device comprises a processor and a memory, wherein a computer program is stored in the memory, and the processor is used to execute the steps in the method according to any one of claims 1 to 10 by calling the computer program stored in the memory.

14. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps in the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Processing method and apparatus for audio data

    CN106024005A

  • Processing method and device for voice recognition in vehicle and electronic equipment

    CN108022591A

  • Audio processing method and device, computer equipment and storage medium

    CN111986691A

  • Echo cancellation method and device, electronic equipment and computer readable storage medium

    CN113299306A

  • Speech enhancement method and device, equipment and storage medium

    CN113450822A