A method and device for reducing noise in audio data, and a computer-readable storage medium

By combining the strong noise reduction model and the speech protection noise reduction model, the noise estimation probability is used for weighted merge to achieve effective noise reduction and voice protection for voice audio, and the problem of difficulty in hearing speech in the prior art is solved.

CN111627455BActive Publication Date: 2025-05-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010495430.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-03
Publication Date
2025-05-09
Estimated Expiration
2040-06-03

AI Technical Summary

Technical Problem

When the prior art reduces the noise of voice audio, it is easy to suppress both voice and noise at the same time, making it difficult to hear the voice clearly.

Method used

An audio data noise reduction method is adopted, and two noise reduction models are obtained through two noise reduction models: a strong noise reduction model and a voice protection noise reduction model. The noise reduction gain is obtained separately, and the noise reduction gain is weighted and combined through the noise estimation probability, and the combined noise reduction gain is determined, and the audio data is noise reduction processed.

Benefits of technology

While maintaining effective noise reduction, try to protect the voice, improve the sound clearness of the voice, and improve the noise reduction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111627455B_ABST
    Figure CN111627455B_ABST
Patent Text Reader

Abstract

The present application discloses an audio data denoising method, device and computer-readable storage medium, the method comprising: obtaining communication audio data; obtaining a first denoising gain for the communication audio data according to a first denoising model, and obtaining a second denoising gain for the communication audio data according to a second denoising model; the denoising intensity of the first denoising model is greater than the denoising intensity of the second denoising model; the degree of voice damage to the communication audio data by the first denoising model is greater than the degree of voice damage to the communication audio data by the second denoising model; determining a combined denoising gain for the communication audio data according to the first denoising gain and the second denoising gain; performing denoising on the communication audio data according to the combined denoising gain to obtain denoised audio data of the communication audio data. The present application can improve the denoising effect on the communication audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio data processing, and in particular to an audio data noise reduction method, device and computer-readable storage medium. Background Art

[0002] In today's society, users usually communicate with each other through the Internet, such as Internet voice communication or Internet phone communication. When user A communicates with user B through the Internet, user A is likely to be in a noisy environment, which causes the voice audio of user A obtained by the terminal held by user A to include not only the voice of user A, but also noise, such as the sound of a car horn, the sound of a TV, or the sound of firecrackers. Therefore, if the unprocessed voice audio of user A is directly sent to user B, it will be difficult for user B to hear the voice of user A clearly. Therefore, it is necessary to perform noise reduction processing on the voice audio recorded by the user.

[0003] In the prior art, when the noise in the voice audio is stronger than the voice, the voice audio will be de-noised to a greater extent, thereby suppressing the noise in the voice audio as much as possible. However, when the noise in the voice audio is suppressed to a greater extent, the voice in the voice audio is usually also suppressed to a greater extent, which makes it difficult for user B to hear the voice of user A based on the de-noised voice audio. As can be seen from the above, the prior art has a poor de-noising effect on the voice audio when de-noising the voice audio. Summary of the invention

[0004] The present application provides an audio data noise reduction method, device and computer-readable storage medium, which can improve the noise reduction effect on communication audio data.

[0005] On one hand, the present application provides an audio data noise reduction method, comprising:

[0006] Acquire communication audio data;

[0007] A first noise reduction gain for the communication audio data is obtained according to the first noise reduction model, and a second noise reduction gain for the communication audio data is obtained according to the second noise reduction model; the noise reduction strength of the first noise reduction model is greater than the noise reduction strength of the second noise reduction model; the degree of voice damage to the communication audio data by the first noise reduction model is greater than the degree of voice damage to the communication audio data by the second noise reduction model;

[0008] Determining a combined noise reduction gain for the communication audio data according to the first noise reduction gain and the second noise reduction gain;

[0009] The communication audio data is subjected to noise reduction processing according to the combined noise reduction gain to obtain noise-reduced audio data of the communication audio data.

[0010] The method of obtaining a first noise reduction gain for the communication audio data according to the first noise reduction model and obtaining a second noise reduction gain for the communication audio data according to the second noise reduction model includes:

[0011] Acquire an audio time domain signal of the communication audio data, and obtain an audio frequency domain signal of the communication audio data according to the audio time domain signal;

[0012] The audio frequency domain signal is input into the first noise reduction model to obtain a first noise reduction gain, and the audio frequency domain signal is input into the second noise reduction model to obtain a second noise reduction gain.

[0013] Wherein, determining a combined noise reduction gain for communication audio data according to the first noise reduction gain and the second noise reduction gain includes:

[0014] Performing a noise estimation operation on the communication audio data to obtain a speech estimation probability of the communication audio data;

[0015] A combined noise reduction gain for the communication audio data is determined according to the speech estimation probability, the first noise reduction gain, and the second noise reduction gain.

[0016] The method of determining a combined noise reduction gain for communication audio data according to the speech estimation probability, the first noise reduction gain, and the second noise reduction gain includes:

[0017] generating a noise weighting coefficient corresponding to a speech estimation probability;

[0018] Weighting the first noise reduction gain according to the noise weighting coefficient to obtain a noise weighted gain;

[0019] Weighting the second noise reduction gain according to the speech estimation probability to obtain a speech weighted gain;

[0020] A combined noise reduction gain is determined according to the noise weighted gain and the speech weighted gain.

[0021] The noise weighting coefficient includes weighting coefficients corresponding to at least two frequency points respectively; the first noise reduction gain includes noise reduction gains corresponding to at least two frequency points respectively; the noise reduction gains corresponding to at least two frequency points included in the first noise reduction gain correspond to the weighting coefficients corresponding to at least two frequency points respectively;

[0022] The first noise reduction gain is weighted according to the noise weighting coefficient to obtain a noise weighted gain, including:

[0023] According to the weighting coefficients corresponding to at least two frequency points in the noise weighting coefficient, the noise reduction gains belonging to the same frequency point in the first noise reduction gain are weighted respectively to obtain the first weighted gain corresponding to each frequency point;

[0024] The first weighted gain corresponding to each frequency point is determined as the noise weighted gain.

[0025] The speech estimation probability includes speech probabilities corresponding to at least two frequency points respectively; the second noise reduction gain includes noise reduction gains corresponding to at least two frequency points respectively; the noise reduction gains corresponding to at least two frequency points included in the second noise reduction gain correspond to the speech probabilities corresponding to at least two frequency points respectively in a one-to-one manner;

[0026] The second noise reduction gain is weighted according to the speech estimation probability to obtain a speech weighted gain, including:

[0027] According to the weighting coefficients corresponding to at least two frequency points in the speech estimation probability, the noise reduction gains belonging to the same frequency point in the second noise reduction gains are weighted respectively to obtain second weighted gains corresponding to each frequency point;

[0028] The second weighted gain corresponding to each frequency point is determined as the speech weighted gain.

[0029] The combined noise reduction gain includes noise reduction gains corresponding to at least two frequency points respectively; the audio frequency domain signal of the communication audio data includes energy values ​​corresponding to at least two frequency points respectively; the noise reduction gains corresponding to at least two frequency points included in the combined noise reduction gain correspond to the energy values ​​corresponding to at least two frequency points respectively in a one-to-one manner;

[0030] The communication audio data is subjected to noise reduction processing according to the combined noise reduction gain to obtain noise-reduced audio data of the communication audio data, including:

[0031] According to the noise reduction gains corresponding to at least two frequency points in the combined noise reduction gain, respectively, energy values ​​belonging to the same frequency point in the communication audio data are weighted to obtain weighted energy values ​​corresponding to each frequency point;

[0032] Determine a weighted audio frequency domain signal of the communication audio data according to the weighted energy value corresponding to each frequency point;

[0033] The weighted audio frequency domain signal is transformed into a time domain to obtain noise reduction audio data of the communication audio data.

[0034] Among them, it also includes:

[0035] Acquire pure speech sample audio data and pure noise sample audio data; the sample audio frequency domain signal of the pure speech sample audio data includes a sample speech energy value; the sample audio frequency domain signal of the pure noise sample audio data includes a sample noise energy value;

[0036] According to the sample speech energy value and the sample noise energy value, the actual noise reduction gain of the sample corresponding to the pure speech sample audio data and the pure noise sample audio data is obtained;

[0037] The sample actual noise reduction gain, pure speech sample audio data and pure noise sample audio data are synchronously input into a first initial noise reduction model, and based on the first initial noise reduction model, a first sample predicted noise reduction gain corresponding to the pure speech sample audio data and the pure noise sample audio data is predicted;

[0038] Based on the actual noise reduction gain of the sample, the first sample predicted noise reduction gain and the first cost function, the model parameters of the first initial noise reduction model are adjusted to obtain the first noise reduction model; the first cost function is used to make the first sample predicted noise reduction gain predicted by the first initial noise reduction model approach the square term of the actual noise reduction gain of the sample.

[0039] Among them, it also includes:

[0040] Acquire pure speech sample audio data and pure noise sample audio data; the sample audio frequency domain signal of the pure speech sample audio data includes a sample speech energy value; the sample audio frequency domain signal of the pure noise sample audio data includes a sample noise energy value;

[0041] According to the sample speech energy value and the sample noise energy value, the actual noise reduction gain of the sample corresponding to the pure speech sample audio data and the pure noise sample audio data is obtained;

[0042] The sample actual noise reduction gain, pure speech sample audio data and pure noise sample audio data are synchronously input into a second initial noise reduction model, and a second sample predicted noise reduction gain corresponding to the pure speech sample audio data and the pure noise sample audio data is predicted based on the second initial noise reduction model;

[0043] Based on the actual noise reduction gain of the sample, the predicted noise reduction gain of the second sample and the second cost function, the model parameters of the second initial noise reduction model are adjusted to obtain the second noise reduction model; the second cost function is used to make the second sample predicted noise reduction gain predicted by the second initial noise reduction model approach the actual noise reduction gain of the sample.

[0044] Among them, it also includes:

[0045] The noise reduction audio data of the communication audio data is synchronized to the connected conversation terminal so that the conversation terminal outputs the noise reduction audio data.

[0046] On one hand, the present application provides an audio data noise reduction device, comprising:

[0047] An audio acquisition module, used to acquire communication audio data;

[0048] A gain acquisition module, used to acquire a first noise reduction gain for communication audio data according to a first noise reduction model, and to acquire a second noise reduction gain for communication audio data according to a second noise reduction model; the noise reduction strength of the first noise reduction model is greater than the noise reduction strength of the second noise reduction model; the degree of voice damage to the communication audio data by the first noise reduction model is greater than the degree of voice damage to the communication audio data by the second noise reduction model;

[0049] A gain combining module, configured to determine a combined noise reduction gain for communication audio data according to the first noise reduction gain and the second noise reduction gain;

[0050] The noise reduction module is used to perform noise reduction processing on the communication audio data according to the combined noise reduction gain to obtain noise-reduced audio data of the communication audio data.

[0051] The gain acquisition module includes:

[0052] A frequency domain signal acquisition unit, used to acquire an audio time domain signal of the communication audio data, and obtain an audio frequency domain signal of the communication audio data according to the audio time domain signal;

[0053] The gain acquisition unit is used to input the audio frequency domain signal into the first noise reduction model to obtain a first noise reduction gain, and input the audio frequency domain signal into the second noise reduction model to obtain a second noise reduction gain.

[0054] Wherein, the gain combining module comprises:

[0055] A noise estimation unit, used to perform a noise estimation operation on the communication audio data to obtain a speech estimation probability of the communication audio data;

[0056] The gain combining unit is used to determine a combined noise reduction gain for the communication audio data according to the speech estimation probability, the first noise reduction gain and the second noise reduction gain.

[0057] Wherein, the gain combining unit comprises:

[0058] A coefficient generating subunit, used for generating a noise weighting coefficient corresponding to the speech estimation probability;

[0059] A first weighting subunit, configured to weight the first noise reduction gain according to the noise weighting coefficient to obtain a noise weighted gain;

[0060] A second weighting subunit, used for weighting the second noise reduction gain according to the speech estimation probability to obtain a speech weighted gain;

[0061] The gain determination subunit is used to determine the combined noise reduction gain according to the noise weighted gain and the speech weighted gain.

[0062] The noise weighting coefficient includes weighting coefficients corresponding to at least two frequency points respectively; the first noise reduction gain includes noise reduction gains corresponding to at least two frequency points respectively; the noise reduction gains corresponding to at least two frequency points included in the first noise reduction gain correspond to the weighting coefficients corresponding to at least two frequency points respectively;

[0063] The first weighted subunit comprises:

[0064] A first frequency point weighting subunit is used to weight the noise reduction gains belonging to the same frequency point in the first noise reduction gains according to the weighting coefficients corresponding to at least two frequency points in the noise weighting coefficient, so as to obtain the first weighted gain corresponding to each frequency point;

[0065] The first frequency point gain determination subunit is used to determine the first weighted gain corresponding to each frequency point as the noise weighted gain.

[0066] The speech estimation probability includes speech probabilities corresponding to at least two frequency points respectively; the second noise reduction gain includes noise reduction gains corresponding to at least two frequency points respectively; the noise reduction gains corresponding to at least two frequency points included in the second noise reduction gain correspond to the speech probabilities corresponding to at least two frequency points respectively in a one-to-one manner;

[0067] The second weighted subunit comprises:

[0068] A second frequency point weighting subunit is used to weight the noise reduction gains belonging to the same frequency point in the second noise reduction gains according to the weighting coefficients corresponding to the at least two frequency points in the speech estimation probability, so as to obtain the second weighted gain corresponding to each frequency point;

[0069] The second frequency point gain determination subunit is used to determine the second weighted gain corresponding to each frequency point as the speech weighted gain.

[0070] The combined noise reduction gain includes noise reduction gains corresponding to at least two frequency points respectively; the audio frequency domain signal of the communication audio data includes energy values ​​corresponding to at least two frequency points respectively; the noise reduction gains corresponding to at least two frequency points included in the combined noise reduction gain correspond to the energy values ​​corresponding to at least two frequency points respectively in a one-to-one manner;

[0071] Noise reduction module, including:

[0072] An energy value weighting unit, used to weight the energy values ​​belonging to the same frequency point in the communication audio data according to the noise reduction gains corresponding to at least two frequency points in the combined noise reduction gain, so as to obtain a weighted energy value corresponding to each frequency point;

[0073] A weighted signal determination unit, used to determine a weighted audio frequency domain signal of the communication audio data according to the weighted energy value corresponding to each frequency point;

[0074] The domain transformation unit is used to perform time domain transformation on the weighted audio frequency domain signal to obtain noise reduction audio data of the communication audio data.

[0075] The audio data noise reduction device further includes:

[0076] The first sample acquisition module is used to acquire pure speech sample audio data and pure noise sample audio data; the sample audio frequency domain signal of the pure speech sample audio data includes a sample speech energy value; the sample audio frequency domain signal of the pure noise sample audio data includes a sample noise energy value;

[0077] A first actual gain acquisition module, used to obtain sample actual noise reduction gains corresponding to pure speech sample audio data and pure noise sample audio data according to the sample speech energy value and the sample noise energy value;

[0078] A first predicted gain acquisition module is used to synchronously input the sample actual noise reduction gain, pure speech sample audio data and pure noise sample audio data into a first initial noise reduction model, and predict the first sample predicted noise reduction gain corresponding to the pure speech sample audio data and the pure noise sample audio data based on the first initial noise reduction model;

[0079] The first parameter adjustment module is used to adjust the model parameters of the first initial denoising model based on the actual noise reduction gain of the sample, the first sample predicted noise reduction gain and the first cost function to obtain the first noise reduction model; the first cost function is used to make the first sample predicted noise reduction gain predicted by the first initial denoising model approach the square term of the actual noise reduction gain of the sample.

[0080] The audio data noise reduction device further includes:

[0081] The second sample acquisition module is used to acquire pure speech sample audio data and pure noise sample audio data; the sample audio frequency domain signal of the pure speech sample audio data includes a sample speech energy value; the sample audio frequency domain signal of the pure noise sample audio data includes a sample noise energy value;

[0082] A second actual gain acquisition module is used to obtain the sample actual noise reduction gain corresponding to the pure speech sample audio data and the pure noise sample audio data according to the sample speech energy value and the sample noise energy value;

[0083] A second predicted gain acquisition module is used to synchronously input the sample actual noise reduction gain, pure speech sample audio data and pure noise sample audio data into a second initial noise reduction model, and predict the second sample predicted noise reduction gain corresponding to the pure speech sample audio data and the pure noise sample audio data based on the second initial noise reduction model;

[0084] The second parameter adjustment module is used to adjust the model parameters of the second initial denoising model based on the actual denoising gain of the sample, the second sample predicted denoising gain and the second cost function to obtain the second denoising model; the second cost function is used to make the second sample predicted denoising gain predicted by the second initial denoising model approach the actual denoising gain of the sample.

[0085] The audio data noise reduction device is also used for:

[0086] The noise reduction audio data of the communication audio data is synchronized to the connected conversation terminal so that the conversation terminal outputs the noise reduction audio data.

[0087] In one aspect, the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes a method in one aspect of the present application.

[0088] In one aspect, the present application provides a computer-readable storage medium storing a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the method in the above aspect.

[0089] The present application obtains communication audio data; obtains a first noise reduction gain for the communication audio data according to a first noise reduction model, and obtains a second noise reduction gain for the communication audio data according to a second noise reduction model; the noise reduction strength of the first noise reduction model is greater than the noise reduction strength of the second noise reduction model; the degree of voice damage to the communication audio data by the first noise reduction model is greater than the degree of voice damage to the communication audio data by the second noise reduction model; determines a combined noise reduction gain for the communication audio data according to the first noise reduction gain and the second noise reduction gain; performs noise reduction processing on the communication audio data according to the combined noise reduction gain to obtain noise-reduced audio data of the communication audio data. It can be seen from this that the method proposed in the present application can obtain a combined noise reduction gain for the communication audio data through a first noise reduction model with relatively large noise reduction strength and a second noise reduction model with relatively good voice protection capability, and then the communication audio data can be denoised by the combined noise reduction gain, so that the voice audio in the communication audio data can be damaged to a lesser extent while the noise audio in the communication audio data is denoised to a greater extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] In order to more clearly illustrate the technical solutions in the present application or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0091] Figure 1 It is a schematic diagram of a system architecture provided by this application;

[0092] Figure 2 This is a schematic diagram of an audio noise reduction scenario provided by this application;

[0093] Figure 3 It is a flowchart of an audio data noise reduction method provided by the present application;

[0094] Figure 4 This is a schematic diagram of a model training scenario provided by this application;

[0095] Figure 5 This is a schematic diagram of a scenario for obtaining a combined noise reduction gain provided by the present application;

[0096] Figure 6 This is a schematic diagram of an audio noise reduction scenario provided by this application;

[0097] Figure 7 This is a schematic diagram of a scenario of an audio noise reduction application provided by this application;

[0098] Figure 8 It is a flowchart of an audio noise reduction method provided by the present application;

[0099] Fig. 9 It is a tabular schematic diagram of experimental data provided by this application;

[0100] Fig.10 It is a structural schematic diagram of an audio data noise reduction device provided by the present application;

[0101] Fig.11 It is a structural schematic diagram of a computer device provided by this application. DETAILED DESCRIPTION

[0102] The following will be combined with the drawings in this application to clearly and completely describe the technical solutions in this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0103] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0104] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0105] This application mainly involves machine learning in artificial intelligence. Among them, machine learning (ML) is a multi-field interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching.

[0106] The machine learning involved in this application mainly refers to obtaining a noise reduction model through machine learning, through which noise reduction of communication audio data can be achieved.

[0107] See also Figure 1 , is a schematic diagram of a system architecture provided by this application. Figure 1As shown, the system architecture diagram includes a server 100 and multiple terminal devices, and the multiple terminal devices specifically include terminal devices 200a, terminal devices 200b and terminal devices 200c. Among them, terminal devices 200a, terminal devices 200b and terminal devices 200c can communicate with each other through the network with the server 100, and the terminal devices can be mobile phones, tablet computers, laptops, PDAs, mobile Internet devices (mobile internet devices, MID), wearable devices (such as smart watches, smart bracelets, etc.). Here, the communication between terminal devices 200a, terminal devices 200b and server 100 is used as an example to illustrate this application.

[0108] See also Figure 2 , is a schematic diagram of an audio noise reduction scenario provided by this application. Figure 2 As shown, the terminal device held by user A may be terminal device 200a, and the terminal device held by user B may be terminal device 200b. User A can use terminal device 200a to make a voice call with user B, and user B can use terminal device 200b to make a voice call with user A. Among them, user A can use the voice call function in the instant messaging application in terminal device 200a to make a voice call with user B, and similarly, user B can use the voice call function in the instant messaging application in terminal device 200b to make a voice call with user A. During the voice call between user A and user B, terminal device 200a can obtain the voice audio recorded by user A, which is the audio of user A speaking to user B. Similarly, terminal device 200b can also obtain the voice audio recorded by user B, which is the audio of user B speaking to user A.

[0109] Among them, in the actual voice call scenario, user A and user B are likely to be in a noisy environment. Therefore, the voice audio of user A obtained by the above terminal device 200a includes not only the voice of user A, but also noise. Similarly, the voice audio of user B obtained by the above terminal device 200b includes not only the voice of user B, but also noise. Assuming that user A is on the street, the noise in the voice audio of user A obtained by the terminal device 200a can be the sound of vehicle horns, the sound of music played by roadside stores, and the voices of passers-by. Assuming that user B is in a shopping mall, the noise in the voice audio of user B obtained by the terminal device 200b can be the sound of music played in the shopping mall, the voices of people in the shopping mall, and the shouting of merchants.

[0110] If the voice audio of user A acquired by terminal device 200a is directly sent to terminal device 200b held by user B, it will be difficult for user B to hear clearly what user A said through the voice audio of user A acquired by terminal device 200b. Therefore, before the voice audio of user A acquired by terminal device 200a is sent to terminal device 200b held by user B, it is necessary to reduce the noise of user A's voice audio, and then send the reduced-noise voice audio of user A to terminal device 200b held by user B, so that user B can quickly hear clearly what user A said through the reduced-noise voice audio of user A acquired by terminal device 200b. Similarly, before the voice audio of user B acquired by terminal device 200b is sent to terminal device 200a held by user A, the voice audio of user B is also reduced in noise, and then the reduced-noise voice audio of user B is sent to terminal device 200a held by user A, so that user A can hear clearly what user A said through the reduced-noise voice audio of user B acquired by terminal device 200a.

[0111] Among them, the noise reduction of voice audio means that the noise in the voice audio needs to be suppressed, and the voice of the user in the voice audio needs to be protected. In other words, the noise reduction of voice audio means reducing the sound of the noise in the voice audio, but trying not to reduce the voice of the user in the voice audio. The following is an example of the noise reduction of the voice audio of user A obtained by the terminal device 200a. It can be understood that the process of noise reduction of the voice audio of user B obtained by the terminal device 200b is the same as the process of noise reduction of the voice audio of user A.

[0112] Among them, when the voice audio of user A obtained by the terminal device 200a is subjected to noise reduction, the terminal device 200a may send the obtained original voice audio of user A to the server 100, and the server 100 performs noise reduction on the voice audio of user A, and then the server 100 may send the noise-reduced voice audio of user A to the terminal device 200b held by user B. Alternatively, after obtaining the voice audio of user A, the terminal device 200a may perform noise reduction on the voice audio of user A by itself, and after noise reduction, send the noise-reduced voice audio of user A to the server 100, and then the server 100 forwards the noise-reduced voice audio of user A to the terminal device 200b held by user B. In other words, the execution subject of the noise reduction of the audio data may be a terminal device or a server, which is determined according to the actual application scenario and is not limited to this. Here, the server is taken as an example to illustrate the execution subject of the noise reduction of the voice audio, and the following describes the process of the server 100 performing noise reduction on the voice audio of user A:

[0113] like Figure 2As shown, it is assumed that the voice audio of user A obtained by the terminal device 200a is noisy communication audio 102a. The terminal device 200a can send the obtained noisy communication audio 102a to the server 100. After the server 100 obtains the noisy communication audio 102a, it can perform a domain transformation on the noisy communication audio 102a, that is, transform the noisy communication audio 102a into the frequency domain, and obtain an audio frequency domain signal of the noisy communication audio 102a. The audio frequency domain signal is a sequence containing multiple energy values ​​(the unit of the energy value is: decibel, i.e. dB), which can be Figure 2 The audio frequency domain signal 107a in the sequence. One energy value in the sequence corresponds to one frequency point, and one frequency point is one frequency sampling point.

[0114] Next, the server 100 can input the audio frequency domain signal of the noisy communication audio 102a into the strong noise reduction model 103a to obtain the noise reduction gain for the noisy communication audio 102a. The audio frequency domain signal of the noisy communication audio 102a can also be input into the voice protection noise reduction model 104a to obtain the noise reduction gain for the noisy communication audio 102a. Among them, the strong noise reduction model 103a and the voice protection noise reduction model are both pre-trained noise reduction models with the ability to reduce noise on audio. The noise reduction capability of the strong noise reduction model 103a is greater than that of the voice protection noise reduction model, and the degree of damage to the voice in the audio by the voice protection noise reduction model is less than that of the strong noise reduction model. In other words, the strong noise reduction model 103a has a greater ability to suppress noise in the audio than the voice protection noise reduction model 104a, and the voice protection noise reduction model 104a has a greater ability to protect the voice in the audio than the strong noise reduction model. The training process of the strong noise reduction model 103a and the voice-protection noise reduction model 104a may refer to the process described in the following step S104.

[0115] The server 100 can also perform noise estimation on the audio frequency domain signal of the noisy communication audio 102a to obtain a speech estimation probability for the audio frequency domain signal of the noisy communication audio 102a. The speech estimation probability indicates the probability that the noisy communication audio 102a is the user's speech rather than noise. After obtaining the speech estimation probability for the noisy communication audio 102a, the server 100 can calculate the final noise reduction gain for the noisy communication audio 102a based on the speech estimation probability 105a, the noise reduction gain for the noisy communication audio 102a obtained by the above-mentioned strong noise reduction model 103a, and the noise reduction gain for the noisy communication audio 102a obtained by the above-mentioned voice protection noise reduction model 104a. The final noise reduction gain is also a gain sequence. Here, the final noise reduction gain can be Figure 2The process of calculating the final noise reduction gain for the noisy communication audio 102a through the speech estimation probability, the noise reduction gain obtained by the strong noise reduction model 103a and the noise reduction gain obtained by the speech protection model can refer to the process described in the following steps S102 and S103.

[0116] like Figure 2 As shown, the gain sequence 106a includes noise reduction gains corresponding to five frequency points, including a noise reduction gain 5 corresponding to frequency point 1, a noise reduction gain 7 corresponding to frequency point 2, a noise reduction gain 8 corresponding to frequency point 3, a noise reduction gain 10 corresponding to frequency point 4, and a noise reduction gain 3 corresponding to frequency point 5. The audio frequency domain signal 107a of the noisy communication audio 102a also includes energy values ​​corresponding to the above five frequency points, specifically including an energy value 1 corresponding to frequency point 1, an energy value 2 corresponding to frequency point 2, an energy value 3 corresponding to frequency point 3, an energy value 2 corresponding to frequency point 4, and an energy value 1 corresponding to frequency point 5.

[0117] The server 100 can achieve noise reduction for the noisy communication audio 102a through the gain sequence 106a: the server 100 can calculate the product between the noise reduction gain and the energy value corresponding to the same frequency point in the gain sequence 106a and the audio frequency domain signal 107a, and obtain the weighted frequency domain signal 108a through the product. Specifically: the server 100 can calculate the product between the noise reduction gain 5 corresponding to frequency point 1 in the gain sequence 106a and the energy value 1 corresponding to frequency point 1 in the audio frequency domain signal 107a, and obtain the weighted energy value, which is the energy value 5 corresponding to frequency point 1 in the weighted frequency domain signal 108a. The server 100 can calculate the product between the noise reduction gain 7 corresponding to frequency point 2 in the gain sequence 106a and the energy value 2 corresponding to frequency point 2 in the audio frequency domain signal 107a, and obtain the weighted energy value, which is the energy value 14 corresponding to frequency point 2 in the weighted frequency domain signal 108a. The server 100 may calculate the product between the noise reduction gain 8 corresponding to the frequency point 3 in the gain sequence 106a and the energy value 3 corresponding to the frequency point 3 in the audio frequency domain signal 107a to obtain a weighted energy value, which is the energy value 24 corresponding to the frequency point 3 in the weighted frequency domain signal 108a. The server 100 may calculate the product between the noise reduction gain 10 corresponding to the frequency point 4 in the gain sequence 106a and the energy value 2 corresponding to the frequency point 4 in the audio frequency domain signal 107a to obtain a weighted energy value, which is the energy value 20 corresponding to the frequency point 4 in the weighted frequency domain signal 108a. The server 100 may calculate the product between the noise reduction gain 3 corresponding to the frequency point 5 in the gain sequence 106a and the energy value 1 corresponding to the frequency point 5 in the audio frequency domain signal 107a to obtain a weighted energy value, which is the energy value 3 corresponding to the frequency point 5 in the weighted frequency domain signal 108a.

[0118] After obtaining the weighted frequency domain signal 108a corresponding to the audio frequency domain signal 107a, the server 100 can perform a time domain transformation on the weighted frequency domain signal 109a to obtain the noise reduction communication audio 109a of the noisy communication audio 102a. The noise reduction communication audio 109a is the final audio obtained after the noise reduction of the noisy communication audio 102a. The server 100 can send the noise reduction communication audio 109a to the terminal device 200b held by user B, and the terminal device 200b can play the noise reduction communication audio 109a so that user B can hear what user A said.

[0119] By adopting the method provided in the present application, a model with strong noise reduction capability (such as the above-mentioned strong noise reduction model 103a) and a model with strong voice protection capability in the audio (such as the above-mentioned voice protection model 104a) are combined to obtain a final noise reduction gain for the audio, and the final noise reduction gain is used to reduce the noise of the audio, so that the audio obtained by the final noise reduction can suppress the noise in the audio to the greatest extent and protect the voice in the audio to the greatest extent, thereby achieving near-ideal noise reduction for the audio.

[0120] See also Figure 3 , is a flowchart of an audio data noise reduction method provided by the present application, such as Figure 3 As shown, the method may include:

[0121] Step S101, acquiring communication audio data;

[0122] Specifically, the execution subject in this embodiment can be a terminal device or a server. If the execution subject is a server, the communication audio data acquired by the server can be sent to it by the terminal device. Here, the target terminal device is used as an example to illustrate the execution subject in this embodiment, and the target terminal device can be any terminal device.

[0123] The target terminal device may obtain the communication audio data from: assuming that the target user holds the target terminal device, the target user can use the target terminal device to make voice calls with other users, and other users also make voice calls with the target user through their own terminal devices. When the target terminal device responds to the user operation of the target user to make a voice call with the terminal device of other users, the target terminal device can obtain the voice call audio recorded by the target user. The voice call audio is what the target user said to other users obtained by the target terminal device, and the voice call audio can be used as the above-mentioned communication audio data.

[0124] In addition, the target terminal device can also respond to the user operation of the target user and conduct an online voice conference with other users. The online voice conference can be a pure voice conference or an online video conference. During the period when the target terminal device is in an online voice conference, the target terminal device can obtain the voice conference audio of the target user. The voice conference audio is what the target user says to other users in the conference during the online voice, and the voice conference audio can also be used as the above-mentioned communication audio data.

[0125] Among them, the target terminal device has communication type software installed, and the target terminal device can realize voice calls between the above-mentioned target user and other users through the voice call function in the installed communication type software; or, the target terminal device can also realize online voice conferences between the above-mentioned target user and other users through the online voice conference function in the installed communication type software.

[0126] It can be understood that the above are only several example scenarios in which the target terminal device obtains communication audio data, and the communication audio data can be any audio data containing user voice obtained in a voice communication related scenario.

[0127] Step S102, obtaining a first noise reduction gain for the communication audio data according to the first noise reduction model, and obtaining a second noise reduction gain for the communication audio data according to the second noise reduction model; the noise reduction strength of the first noise reduction model is greater than the noise reduction strength of the second noise reduction model; the degree of voice damage to the communication audio data by the first noise reduction model is greater than the degree of voice damage to the communication audio data by the second noise reduction model;

[0128] Specifically, when the target user uses the target terminal device to conduct the above-mentioned voice call or online voice conference with other users, it is very likely that the target user is in a noisy environment. Therefore, the communication audio data obtained by the target terminal device will contain not only the voice of the target user, but also noise. For example, if the target user is in an air-conditioned room, the noise in the communication audio data can be the running sound of the air conditioner or the rotation sound of the electric fan, etc.; if the target user is in a shopping mall, the noise in the communication audio data can be the sound of the music played in the shopping mall and the shouting of the clerk, etc.; if the target user is on the street, the noise in the communication audio data can be the horn of the vehicle and the voice of the passerby. Therefore, this embodiment mainly describes how to reduce the noise of the communication audio data, and the effect that the noise reduction wants to achieve is to suppress the noise in the communication audio data as much as possible, and protect the voice of the user in the communication audio data as much as possible.

[0129] The first noise reduction model (which may be the above Figure 2 The strong noise reduction model in the above) and the second noise reduction model (which can be the above Figure 2The first and second noise reduction models are pre-trained models that can obtain noise reduction gains for communication audio data. In other words, the first noise reduction model and the second noise reduction model can both achieve noise reduction for communication audio data. Among them, the noise reduction intensity of the first noise reduction model for communication audio data is greater than the noise reduction intensity of the second noise reduction model for communication audio data. It can be understood that the greater the noise reduction intensity, although the suppression intensity of the noise in the communication audio data is greater, it is also easier to cause damage to the user's voice in the communication audio data, for example, the user's voice in the communication audio data is also suppressed too much. Although the noise reduction intensity of the second noise reduction model for communication audio data is less than that of the first noise reduction model, the degree of damage to the user's voice in the communication audio data by the second noise reduction model is less than that of the first noise reduction model. In other words, the degree of protection of the user's voice in the communication audio data by the second noise reduction model is greater than that of the first noise reduction model.

[0130] The following describes the training process of the first denoising model and the second denoising model:

[0131] Sample audio data may be acquired in advance, and the sample audio data may include pure voice sample audio data and pure noise sample audio data. The pure voice sample audio data includes only the sound of the user speaking, and the pure voice sample audio data may be the sound of various users speaking, such as the sound of users speaking with different voices, or the sound of users speaking with different timbres. The pure noise sample audio data includes only noise, for example, the pure noise sample audio data may be various types of noise such as the sound of vehicle horns, the sound of cooking, or the sound of keyboard tapping. The pure voice sample audio data and the pure noise sample audio data are input into the initial model for training to obtain the above-mentioned first noise reduction model and the second noise reduction model.

[0132] It can be understood that a pure speech sample audio data and a pure noise sample audio data constitute a sample. Among them, the pure speech sample audio data can be transformed in the time domain to obtain the time domain signal of the pure speech sample audio data, and then the time domain signal of the pure speech sample audio data can be transformed in the frequency domain to obtain the frequency domain signal of the pure speech sample audio data. Similarly, the pure noise sample audio data can be transformed in the time domain to obtain the time domain signal of the pure noise sample audio data, and then the time domain signal of the pure noise sample audio data can be transformed in the frequency domain to obtain the frequency domain signal of the pure noise sample audio data.

[0133] Among them, the frequency domain signal of the pure speech sample audio data and the frequency domain signal of the pure noise sample audio data belonging to a sample have the same signal length. For example, the frequency domain signal of the pure speech sample audio data belonging to a sample contains energy values ​​corresponding to 3 frequency points, such as (1, 2, 3), then the frequency domain signal of the pure noise sample audio data in the sample also contains energy values ​​corresponding to 3 frequency points, such as (4, 5, 6). The energy value in the frequency domain signal of the pure speech sample audio data and the energy value in the frequency domain signal of the pure noise sample audio data both correspond to each frequency point one by one. The frequency domain signal of the pure speech sample audio data and the frequency domain signal of the pure noise sample audio data can be referred to as sample audio frequency domain signals. The energy value in the frequency domain signal of the pure speech sample audio data can be referred to as a sample speech energy value, and the energy value in the frequency domain signal of the pure noise sample audio data can be referred to as a sample noise energy value. Among them, a frequency point is a frequency sampling point. Assuming that the above three frequency points include frequency point 1, frequency point 2 and frequency point 3, then the energy value 1 in the frequency domain signal (1, 2, 3) of the above pure speech sample audio data and the energy value 4 in the frequency domain signal (4, 5, 6) of the pure noise sample audio data both correspond to frequency point 1; the energy value 2 in the frequency domain signal (1, 2, 3) of the pure speech sample audio data and the energy value 5 in the frequency domain signal (4, 5, 6) of the pure noise sample audio data both correspond to frequency point 2; the energy value 3 in the frequency domain signal (1, 2, 3) of the pure speech sample audio data and the energy value 6 in the frequency domain signal (4, 5, 6) of the pure noise sample audio data both correspond to frequency point 3.

[0134] The ratio between the energy values ​​corresponding to the same frequency point in the frequency domain signal of the pure speech sample audio data and the frequency domain signal of the pure noise sample audio data can be calculated to obtain the actual noise reduction gain corresponding to each frequency point of the sample. The actual noise reduction gain corresponding to each frequency point of the sample can be called the actual noise reduction gain of the sample. For example, a ratio of 1 / 3 between the energy value 1 in the frequency domain signal (1, 2, 3) of the pure speech sample audio data and the energy value 4 in the frequency domain signal (4, 5, 6) of the pure noise sample audio data can be obtained, and the ratio 1 / 3 is the actual noise reduction gain of the sample corresponding to frequency point 1; a ratio of 2 / 5 between the energy value 2 in the frequency domain signal (1, 2, 3) of the pure speech sample audio data and the energy value 5 in the frequency domain signal (4, 5, 6) of the pure noise sample audio data can be obtained, and the ratio 2 / 5 is the actual noise reduction gain of the sample corresponding to frequency point 2; a ratio of 3 / 6 between the energy value 3 in the frequency domain signal (1, 2, 3) of the pure speech sample audio data and the energy value 6 in the frequency domain signal (4, 5, 6) of the pure noise sample audio data can be obtained, and the ratio 3 / 6 is the actual noise reduction gain of the sample corresponding to frequency point 3.

[0135] The training process of the first denoising model is described here:

[0136] A sample may include a pure speech sample audio data and a pure noise sample audio data, and multiple samples and the actual noise reduction gain of each sample (which may be a gain sequence, including the actual noise reduction gain of each frequency point) may be input into the first initial noise reduction model for training, wherein the model structure of the first initial noise reduction model may be a DNN (deep neural network) network structure. The cost function of the first initial noise reduction model may be referred to as a first cost function, and the first cost function is shown in formula (1):

[0137]

[0138]

[0139] It can be seen that the first cost function L1 is the square term of the MSE (mean square error) criterion, n is the frequency point, the initial value of n is 1, and there are n frequency points in total. is the predicted noise reduction gain corresponding to the nth frequency point predicted by the first initial noise reduction model. The predicted noise reduction gain corresponding to each frequency point predicted by the first initial noise reduction model can be referred to as the first sample predicted noise reduction gain. n is the actual noise reduction gain of the sample corresponding to the nth frequency point. 1n is the energy value corresponding to the frequency point n of the pure speech sample audio data, E 2n is the energy value corresponding to the frequency point n of the pure noise sample audio data.

[0140] When training the first initial denoising model, the model parameters of the first initial denoising model are adjusted so that the first cost function reaches the minimum value, that is, the training loss is minimized. By training the first initial denoising model through the first cost function, the first sample predicted denoising gain predicted by the first initial denoising model can be infinitely close to the square term of the sample actual denoising gain. When the first initial denoising model is trained to convergence, the converged first initial denoising model can be used as the above-mentioned first denoising model.

[0141] The training process of the second denoising model is described here:

[0142] The second noise reduction model can be trained using the same sample data as the first noise reduction model, or it can be trained using different sample data. When training the second initial noise reduction model to obtain the second noise reduction model, the biggest difference from training the first initial noise reduction model to obtain the second noise reduction model is that the cost function of the second initial noise reduction model is different from the cost function of the first initial noise reduction model. The network structure of the second initial noise reduction model can also be a DNN network structure. Similarly, a sample can include a pure speech sample audio data and a pure noise sample audio data, and multiple samples and the actual noise reduction gain of each sample corresponding to the sample can be input into the second initial noise reduction model for training. The cost function of the second initial noise reduction model can be called the second cost function, and the second cost function is shown in formula (2):

[0143]

[0144] in, represents the MSE criterion term, represents the traditional cross entropy cost function (CrossEntropy Loss). n represents the frequency point. is the predicted noise reduction gain corresponding to the nth frequency point predicted by the second initial noise reduction model. The predicted noise reduction gain corresponding to each frequency point predicted by the second initial noise reduction model can be referred to as the second sample predicted noise reduction gain. n It is the actual noise reduction gain of the sample corresponding to the nth frequency point.

[0145] When training the second initial denoising model, the model parameters of the second initial denoising model are adjusted so that the second cost function reaches the minimum value, that is, the training loss is minimized. By training the second initial denoising model with the second cost function, the second sample predicted denoising gain predicted by the second initial denoising model can be infinitely close to the actual denoising gain of the sample. When the first initial denoising model is trained to convergence, the converged first initial denoising model can be used as the above-mentioned first denoising model.

[0146] Among them, it can be understood that the main function of the above-mentioned MSE criterion item is noise reduction (i.e., suppressing the noise in the communication audio data), and the main function of the traditional cross entropy cost function L is to protect the voice in the communication audio data. By using different cost functions to train the initial model, the direction of the initial model training can be limited. The above-mentioned training of the first initial noise reduction model by the first cost function is to make the training direction of the first initial noise reduction model mainly focus on noise reduction, and the training of the second initial noise reduction model by the above-mentioned second cost function is to make the training direction of the second initial noise reduction model mainly focus on protecting the voice on the basis of noise reduction.

[0147] Therefore, through the above process, a first noise reduction model with stronger noise reduction capability and a second noise reduction model with stronger voice protection capability can be trained.

[0148] See also Figure 4 , is a schematic diagram of a model training scenario provided by this application. Figure 4 As shown, the sample data 100f used for model training includes pure speech sample audio data 101f and pure noise sample audio data 102f. The pure speech sample audio data 101f is an audio set, which includes pure speech sample audio data 103f (i.e., {5, 4, 3, 2, 1}). The pure speech sample audio data 103f includes an energy value 5 corresponding to frequency point 1, an energy value 4 corresponding to frequency point 2, an energy value 3 corresponding to frequency point 3, an energy value 2 corresponding to frequency point 4, and an energy value 1 corresponding to frequency point 5. The pure noise sample audio data 102f is also an audio set, which includes pure noise sample audio data 104f (i.e., {1, 2, 3, 4, 5}). The pure noise sample audio data 104f includes an energy value 1 corresponding to frequency point 1, an energy value 2 corresponding to frequency point 2, an energy value 3 corresponding to frequency point 3, an energy value 4 corresponding to frequency point 4, and an energy value 5 corresponding to frequency point 5. The pure speech sample audio data 103f and the pure noise sample audio data 104f belong to the same sample, so the sample actual noise reduction gain obtained by the above pure speech sample audio data 103f and the pure noise sample audio data 104f is the sample actual noise reduction gain 105f (i.e. {5 / 1, 4 / 2, 3 / 3, 2 / 4, 1 / 5}, that is, {5, 2, 1, 0.5, 0.2}). The sample actual noise reduction gain 105f includes the actual noise reduction gain 5 corresponding to frequency point 1, the actual noise reduction gain 2 corresponding to frequency point 2, the actual noise reduction gain 1 corresponding to frequency point 3, the actual noise reduction gain 0.5 corresponding to frequency point 4, and the actual noise reduction gain 0.2 corresponding to frequency point 5. Therefore, the sample actual noise reduction gain corresponding to each sample in the sample data 100f can be obtained in this way.

[0149] Therefore, each sample consisting of the pure speech sample audio data 101f and the pure noise sample audio data 102f in the sample data 100f and the actual noise reduction gain of each sample can be input into the first initial noise reduction model 108f for training, wherein the pure speech sample audio data and the pure noise sample audio data belonging to the same sample are synchronously input into the first initial noise reduction model 108f. During the training process, the first initial noise reduction model is trained by the first cost function 106f. After the training of the first initial noise reduction model is completed, the first noise reduction model 110f can be obtained.

[0150] Similarly, each sample consisting of the pure speech sample audio data 101f and the pure noise sample audio data 102f in the sample data 100f and the actual noise reduction gain of each sample corresponding to the sample can be input into the second initial noise reduction model 109f for training, wherein the pure speech sample audio data and the pure noise sample audio data belonging to the same sample are synchronously input into the second initial noise reduction model 109f. During the training process, the second initial noise reduction model is trained by the second cost function 107f. After the training of the second initial noise reduction model is completed, the second noise reduction model 111f can be obtained.

[0151] Among them, the standard for completing model training can be that the error of the final trained model is within a reasonable range (the reasonable range can be set by yourself), or the number of training sample data reaches a certain value (the value can be set by yourself), etc.

[0152] Next, the above-mentioned communication audio data can be transformed in the time domain to obtain the time domain signal of the communication audio data, and the time domain signal of the communication audio data can be transformed in the frequency domain (for example, windowed FFT (Fast Fourier Transform) transformation) to obtain the frequency domain signal of the communication audio data (which can be called the audio frequency domain signal). The frequency domain signal of the communication audio data can be input into the first noise reduction model trained above to obtain the first noise reduction gain. The frequency domain signal of the communication audio data can be input into the second noise reduction model trained above to obtain the second noise reduction gain.

[0153] Among them, the audio frequency domain signal of the communication audio data includes energy values ​​corresponding to at least two frequency points respectively. For example, the audio frequency domain signal of the communication audio data is (1, 2, 3), and the audio frequency domain signal (1, 2, 3) includes energy value 1 corresponding to frequency point 1, energy value 2 corresponding to frequency point 2, and energy value 3 corresponding to frequency point 3. The first noise reduction gain obtained above also includes the noise reduction gains corresponding to the at least two frequency points respectively, and the second noise reduction gain also includes the noise reduction gains corresponding to the at least two frequency points respectively. For example, the audio frequency domain signal of the communication audio data includes energy value 1 corresponding to frequency point 1, energy value 2 corresponding to frequency point 2, and energy value 3 corresponding to frequency point 3, then the first noise reduction gain can include the noise reduction gain corresponding to frequency point 1, the noise reduction gain corresponding to frequency point 2, and the noise reduction gain corresponding to frequency point 3. Similarly, the first noise reduction gain can also include the noise reduction gain corresponding to frequency point 1, the noise reduction gain corresponding to frequency point 2, and the noise reduction gain corresponding to frequency point 3.

[0154] It can be understood that the noise reduction gain in the first noise reduction gain corresponds one-to-one to the energy value in the audio frequency domain signal of the communication audio data, and the noise reduction gain in the second noise reduction gain also corresponds one-to-one to the energy value in the audio frequency domain signal of the communication audio data, and the corresponding method is to correspond to the same frequency point.

[0155] Step S103, determining a combined noise reduction gain for the communication audio data according to the first noise reduction gain and the second noise reduction gain;

[0156] Specifically, the target terminal device can perform noise estimation on the above-mentioned communication audio data to obtain the speech estimation probability of the communication audio data (also referred to as the speech existence probability), wherein the speech estimation probability includes the speech probability corresponding to each energy value in the audio frequency domain signal of the communication audio data, and the speech probability indicates the probability of speech existence at the corresponding energy value. Among them, the method of performing noise estimation on the communication audio data to obtain the speech estimation probability of the communication audio data can be a method of using MCR (minimum controlled recursive averaging) noise estimation, a noise estimation method using speech correlation, or a noise estimation method using noise correlation.

[0157] The above-mentioned speech estimation probability can be denoted as p, the above-mentioned first noise reduction gain can be denoted as gain1, the above-mentioned second noise reduction gain can be denoted as gain2, and the combined noise reduction gain can be denoted as gain. The method of obtaining the combined noise reduction gain by using the above-mentioned first noise reduction gain, the second noise reduction gain and the speech estimation probability can be referred to in the following formula (3):

[0158] gain=(1-p)*gain1+p*gain2 (3)

[0159] Among them, 1-p can be called the noise weighting coefficient corresponding to the speech estimation probability, and the noise weighting coefficient also includes the weighting coefficients corresponding to each frequency point. Through the above process, it can be known that the second noise reduction gain gain2 is weighted by the speech estimation probability p, and the first noise reduction gain is weighted by the noise weighting coefficient 1-p. The result p*gain2 after weighting the second noise reduction gain by the speech estimation probability can be called the speech weighted gain, and the result (1-p)*gain1 after weighting the first noise reduction gain by the noise weighting coefficient can be called the noise weighted gain. By adding the noise weighted gain to the speech weighted gain, the combined noise reduction gain gain can be obtained.

[0160] Since the above-mentioned speech estimation probability, the first noise reduction gain, the second noise reduction gain and the combined noise reduction gain all include the values ​​corresponding to each frequency point, the above-mentioned speech estimation probability can also be recorded as p(k, n), the above-mentioned first noise reduction gain can also be recorded as gain1(k, n), the above-mentioned second noise reduction gain can also be recorded as gain2(k, n), and the above-mentioned combined noise reduction gain can also be recorded as gain(k, n). Wherein, k represents the frequency point, and n represents the number of frames. Since the communication audio data is framed when noise reduction is performed, for example, every 20ms is a frame, etc. Therefore, it is necessary to perform noise reduction on each frame of the communication audio data separately. When n takes the first frame, k can take each frequency point, when n takes the second frame, k can also take each frequency point, when n takes the third frame, k can also take each frequency point, .... It can be understood that the process of denoising a frame of audio of the communication audio data described in this embodiment is the same as that of denoising each frame of audio of the communication audio data. When the noise reduction of all frames of the communication audio data is completed, it indicates that the noise reduction of the communication audio data is completed.

[0161] Therefore, the above formula 1 can also be expressed in the form of the following formula (4):

[0162] gain(k,n)=[1-p(k,n)]*gain1(k,n)+p(k,n)*gain2(k,n) (4)

[0163] It can be seen from the above formula (4) that when the first noise reduction gain is weighted by the noise weighting coefficient, the weighting coefficient and the noise reduction gain corresponding to the same frequency point in the noise weighting coefficient and the first noise reduction gain are weighted (i.e., multiplied). After the first noise reduction gain is weighted by the noise weighting coefficient, the weighted noise reduction gain corresponding to each frequency point is referred to as the first weighted gain. The noise weighted gain includes the first weighted gain corresponding to each frequency point.

[0164] When the second noise reduction gain is weighted by the speech estimation probability, the speech probability corresponding to the same frequency point in the speech estimation probability and the noise reduction gain in the second noise reduction gain are weighted (i.e., multiplied). After the second noise reduction gain is weighted by the speech estimation probability, the weighted noise reduction gain corresponding to each frequency point is referred to as the second weighted gain. The speech weighted gain includes the second weighted gain corresponding to each frequency point.

[0165] Among them, the above is explained with the first noise reduction model and the second noise reduction model being one. Optionally, there may be multiple first noise reduction models, and there may be multiple second noise reduction models. The cost functions for training multiple first noise reduction models may all be the above first cost functions, and the cost functions for training multiple second noise reduction models may all be the above second cost functions. However, when different first noise reduction models are obtained through training, different sample data sets may be used for training, or first initial noise reduction models with different network structures may be used, and one first noise reduction model corresponds to one first initial noise reduction model. Therefore, the noise reduction intensity of each trained first noise reduction model will still be greater than that of the second noise reduction model, but the noise reduction capabilities of different first noise reduction models are still different. Similarly, when different second noise reduction models are obtained through training, different sample data sets may be used for training, or second initial noise reduction models with different network structures may be used, and one second noise reduction model corresponds to one second initial noise reduction model. Therefore, the protection capability of each trained second noise reduction model for speech will still be greater than that of the first noise reduction model, but the protection capability of different second noise reduction models for speech is still different.

[0166] The above-mentioned noise weighting coefficient can be used to weight the first noise reduction gain obtained by each first noise reduction model, and the weighted noise weighted gain corresponding to each first noise reduction model can be summed and averaged, and the average value can be used as the final noise weighted gain corresponding to all first noise reduction models. Similarly, the above-mentioned speech estimation probability can be used to weight the second noise reduction gain obtained by each second noise reduction model, and the weighted speech weighted gain corresponding to each second noise reduction model can be summed and averaged, and the average value can be used as the final speech weighted gain corresponding to all second noise reduction models. Then, the final noise weighted gain corresponding to all first noise reduction models is added to the final speech weighted gain corresponding to all second noise reduction models, and the final combined noise reduction gain for the communication audio data can be obtained. For example, if there is a first noise reduction model m1 and a first noise reduction model m2, there is a second noise reduction model m3 and a second noise reduction model m4. The first noise reduction gain obtained by the first noise reduction model m1 is (1, 2), the first noise reduction gain obtained by the first noise reduction model m2 is (3, 4), the second noise reduction gain obtained by the second noise reduction model m3 is (5, 6), and the second noise reduction gain obtained by the second noise reduction model m4 is (7, 8). The speech estimation probability is (0.2, 0.3), and the noise weighting coefficient is (0.8, 0.7).

[0167] Then, by weighting the first noise reduction gain (1, 2) with the noise weighting coefficient (0.8, 0.7), the noise weighted gain (0.8, 1.4) can be obtained, and by weighting the first noise reduction gain (3, 4) with the noise weighting coefficient (0.8, 0.7), the noise weighted gain (2.4, 2.8) can be obtained. Then, ((0.8+2.4) / 2, (1.4+2.8) / 2), that is, (1.6, 2.1) can be used as the final noise weighted gain. Similarly, by weighting the second noise reduction gain (5, 6) with the speech estimation probability (0.2, 0.3), the speech weighted gain (1.0, 1.8) can be obtained, and by weighting the second noise reduction gain (7, 8) with the speech estimation probability (0.2, 0.3), the speech weighted gain (1.4, 2.4) can be obtained. Then, ((1.0+1.4) / 2, (1.8+2.4) / 2), that is, (1.2, 2.1), can be used as the final speech weighted gain. The above final noise weighted gain (1.6, 2.1) can be added to the final speech weighted gain (1.2, 2.1) to obtain the final combined noise reduction gain (1.6+1.2, 2.1+2.1), that is, (2.8, 4.2).

[0168] See also Figure 5 , is a schematic diagram of a scenario for obtaining a combined noise reduction gain provided by the present application. The model set 100d includes a plurality of first noise reduction models, such as a first noise reduction model 101d, a first noise reduction model 102d, and a first noise reduction model 103d. Each first noise reduction model in the model set 100d can obtain a first noise reduction gain for communication audio data (including a first noise reduction gain 101e obtained by the first noise reduction model 101d, a first noise reduction gain 102e obtained by the first noise reduction model 102d, and a first noise reduction gain 103e obtained by the first noise reduction model 103d). The first noise reduction gain obtained by each first noise reduction model can be weighted using a noise weighting coefficient 104e, thereby obtaining a final noise weighted gain 105e corresponding to all first noise reduction models.

[0169] Similarly, the model set 104d includes a plurality of second noise reduction models such as a second noise reduction model 105d, a second noise reduction model 106d, and a second noise reduction model 107d. Each second noise reduction model in the model set 100d can obtain a second noise reduction gain for the communication audio data (including a second noise reduction gain 106e obtained by the second noise reduction model 105d, a second noise reduction gain 107e obtained by the second noise reduction model 106d, and a second noise reduction gain 108e obtained by the second noise reduction model 107d). The second noise reduction gain obtained by each second noise reduction model can be weighted using the speech estimation probability 109e to obtain the final speech weighted gain 110e corresponding to all second noise reduction models. The above-mentioned noise weighted gain 105e can be added to the speech weighted gain 110e to obtain the final combined noise reduction gain 111e for the communication audio data.

[0170] Step S104, performing noise reduction processing on the communication audio data according to the combined noise reduction gain to obtain noise-reduced audio data of the communication audio data;

[0171] Specifically, the above-mentioned combined noise reduction gain includes the noise reduction gain corresponding to each frequency point, and the audio frequency domain signal of the communication audio data also includes the energy value corresponding to each frequency point. By combining the noise reduction gain corresponding to each frequency point in the noise reduction gain, the energy values ​​belonging to the same frequency point in the audio frequency domain signal of the communication audio data can be weighted (i.e., multiplied) to obtain the weighted energy value corresponding to each frequency point, and the weighted energy value corresponding to each frequency point can be called the weighted energy value.

[0172] The weighted energy values ​​corresponding to each frequency point can be combined to obtain a weighted audio frequency domain signal of the communication audio data, that is, the weighted audio frequency domain signal includes the weighted energy value corresponding to each frequency point. The weighted audio frequency domain signal can be transformed from the frequency domain to the time domain to obtain the final noise reduction audio data of the communication audio data.

[0173] The target terminal device can send the noise reduction audio data of the communication audio data to the connected session terminal, so that the session terminal that obtains the noise reduction audio data can play the noise reduction audio data through the audio player. The session terminal connected to the target terminal device may refer to the terminal device held by the user who has a voice call with the target user in the above step S101, or the terminal device held by the user who has an online voice conference with the target user. In addition, when the target terminal device sends the noise reduction audio data of the communication audio data to the connected session terminal, after the noise reduction of each frame of audio (for example, a frame of audio of 20ms) of the communication audio data is completed, the noise reduction audio data corresponding to the completed frame of audio can be sent to the connected session terminal, without waiting until the noise reduction of the entire communication audio data is completed, and then sending the noise reduction audio data corresponding to the entire communication audio data to the connected session terminal, which can ensure that the connected session terminal can obtain the noise reduction audio data of the call audio data recorded by the target user in real time (with minimal delay).

[0174] In actual application scenarios, when the target user is conducting voice communication with other users, the target user can also obtain the noise-reduced audio data of the communication audio data recorded by other users through the target terminal device. Moreover, the noise-reduced audio data of the communication audio data recorded by other users is also obtained after noise reduction of the communication audio data of other users by the terminal devices held by other users (the noise reduction method is the same as above).

[0175] In the present application, since the combined noise reduction gain for noise reduction processing of communication audio data is obtained by combining the first noise reduction model with relatively large noise reduction strength and the second noise reduction model with relatively large voice protection capability, the combined noise reduction gain finally obtained can achieve the goal of reducing the degree of damage to the voice in the communication audio data and effectively reducing the noise in the communication audio data when noise reduction processing is performed on the communication audio data. Therefore, by adopting the method provided by the present application, a good noise reduction effect can be obtained for the communication audio data.

[0176] See also Figure 6 , is a schematic diagram of an audio noise reduction scenario provided by this application. Figure 6As shown, the audio frequency domain signal of the communication audio data 100b is an audio frequency domain signal 101b (i.e., {1, 2, 3, 4, 5}). The audio frequency domain signal 101b can be input into the first noise reduction model 102b to obtain a first noise reduction gain 104b (i.e., {5, 4, 3, 2, 1}), which includes a noise reduction gain 5 corresponding to frequency point 1, a noise reduction gain 4 corresponding to frequency point 2, a noise reduction gain 3 corresponding to frequency point 3, a noise reduction gain 2 corresponding to frequency point 4, and a noise reduction gain 1 corresponding to frequency point 5. The audio frequency domain signal 101b can be input into the second noise reduction model 103b to obtain a second noise reduction gain 105b (i.e., {9, 7, 5, 3, 1}), which includes a noise reduction gain 9 corresponding to frequency 1, a noise reduction gain 7 corresponding to frequency 2, a noise reduction gain 5 corresponding to frequency 3, a noise reduction gain 3 corresponding to frequency 4, and a noise reduction gain 1 corresponding to frequency 5. The first noise reduction gain 104b (i.e., {5, 4, 3, 2, 1}) can be weighted (weighted for the same frequency) using a noise weighting coefficient 106b (i.e., {0.9, 0.8, 0.7, 0.6, 0.5}) to obtain a noise weighted gain 108b (i.e., {4.5, 3.2, 2.1, 1.2, 0.5}). The second noise reduction gain 105b (i.e., {9, 7, 5, 3, 1}) can be weighted (weighted at the same frequency point) using the speech estimation probability (i.e., {0.1, 0.2, 0.3, 0.4, 0.5}) to obtain a speech weighted gain (i.e., {0.9, 1.4, 1.5, 1.2, 0.5}).

[0177] The noise weighted gain 108b (i.e., {4.5, 3.2, 2.1, 1.2, 0.5}) and the noise reduction gain corresponding to the same frequency point in the speech weighted gain (i.e., {0.9, 1.4, 1.5, 1.2, 0.5}) can be added to obtain a combined noise reduction gain 110b (i.e., {5.4, 4.6, 3.6, 2.4, 1.0}). The audio frequency domain signal 101b of the communication audio signal can be weighted using the combined noise reduction gain 110b (i.e., {5.4, 4.6, 3.6, 2.4, 1.0}) to obtain a weighted audio frequency domain signal 111b (i.e., {5.4, 9.2, 10.8, 9.6, 5}), and then the weighted audio frequency domain signal 111b (i.e., {5.4, 9.2, 10.8, 9.6, 5}) is transformed in the time domain to obtain the noise reduction audio data 112b of the communication audio data 100b.

[0178] See also Figure 7 , is a schematic diagram of a scenario of an audio noise reduction application provided by this application. Figure 7As shown, it can be known from the conference page 101c that the user "Tiantian" has initiated an online conference, and the users "Tiantian", "Lele" and "Duoduo" have all participated in the online conference. The terminal device held by the user "Tiantian" is the terminal device 102c, the terminal device held by the user "Lele" is the terminal device 103c, and the terminal device held by the user "Duoduo" is the terminal device 104c. The terminal device 102c can obtain the communication audio data entered by the user "Tiantian" during the conference, and the terminal device 102c can send the obtained communication audio data of the user "Tiantian" to the server 105c. Similarly, the terminal device 103c can obtain the communication audio data entered by the user "Lele" during the conference, and the terminal device 103c can send the obtained communication audio data of the user "Lele" to the server 105c. Likewise, the terminal device 104c may obtain the communication audio data recorded by the user "Duoduo" during the conference, and the terminal device 104c may send the obtained communication audio data of the user "Duoduo" to the server 105c.

[0179] The server 105c may perform noise reduction on the acquired communication audio data of the user "Tiantian" to obtain the noise-reduced audio data of the communication audio data of the user "Tiantian", perform noise reduction on the acquired communication audio data of the user "Lele" to obtain the noise-reduced audio data of the communication audio data of the user "Lele", and perform noise reduction on the acquired communication audio data of the user "Duoduo" to obtain the noise-reduced audio data of the communication audio data of the user "Duoduo". The server 105c may send the noise-reduced audio data of the user "Lele" and the user "Duoduo" to the user "Tiantian", send the noise-reduced audio data of the user "Tiantian" and the user "Duoduo" to the user "Lele", and send the noise-reduced audio data of the user "Lele" and the user "Tiantian" to the user "Duoduo", so as to realize a conference call between the user "Tiantian", the user "Lele" and the user "Duoduo".

[0180] See also Figure 8 , is a flowchart of an audio noise reduction method provided by this application. Figure 8As shown, step ①: first collect the signal to obtain the above-mentioned communication audio data, and then perform a windowed Fourier transform on the communication audio data to obtain the audio frequency domain signal of the communication audio data. Step ②: input the audio frequency domain signal of the communication audio data into the first noise reduction model to obtain the above-mentioned first noise reduction gain. Step ③: input the audio frequency domain signal of the communication audio data into the second noise reduction model to obtain the second noise reduction gain. Step ④: perform noise estimation on the audio frequency domain signal of the communication audio data through the noise estimation module to obtain the speech estimation probability of the communication audio data. Step ⑤: perform gain fusion on the first noise reduction gain obtained by the first noise reduction model and the second noise reduction gain obtained by the second noise reduction model through the gain fusion module to obtain a combined noise reduction gain. Step ⑥: weight the audio frequency domain signal of the communication audio data by combining the noise reduction gain (i.e., noise reduction processing) to obtain a weighted audio frequency domain signal, and then perform a windowed Fourier inverse transform on the weighted audio frequency domain signal to obtain an output signal, which is the noise reduction audio data of the communication audio data.

[0181] See also Fig. 9 , is a table diagram of experimental data provided by this application. Fig. 9 As shown, the voice quality represents the protection of the voice in the audio during the noise reduction process. The higher the voice quality score, the better the protection of the voice in the audio. The noise reduction quality (also known as the background noise transmission quality) represents the noise reduction of the noise in the audio during the noise reduction process. The higher the noise reduction quality score, the better the noise reduction of the audio. The overall quality is the weighted sum of the voice quality and the noise reduction quality, which represents the overall noise reduction effect on the audio. The higher the overall quality score, the better the overall noise reduction effect on the audio. In actual applications, it is usually hoped that the overall quality score will be higher.

[0182] exist Fig. 9 Table 1 in shows the noise reduction scores for office noise. When the first noise reduction model is used alone to reduce the noise of the office, the overall quality score of the noise reduction is 4.2, the overall quality score of the noise reduction is 4.24 when the second noise reduction model is used alone, and the overall quality score of the noise reduction is 4.31 when the first noise reduction model and the second noise reduction model are used together. Therefore, using the first noise reduction model and the second noise reduction model at the same time has a higher overall quality score for the noise reduction of the office noise and a better noise reduction effect than using the first noise reduction model or the second noise reduction model alone.

[0183] exist Fig. 9Table 2 in shows the noise reduction scores for restaurant noise. When the first noise reduction model is used alone to reduce the noise of the restaurant, the overall quality score of the noise reduction is 3.10, the overall quality score of the noise reduction is 3.02 when the second noise reduction model is used alone, and the overall quality score of the noise reduction is 3.21 when the first noise reduction model and the second noise reduction model are used together. Therefore, the overall quality score of the noise reduction of the restaurant noise is higher and the noise reduction effect is better when the first noise reduction model and the second noise reduction model are used together than when the first noise reduction model or the second noise reduction model is used alone.

[0184] The above experimental data show that the method provided by the present application can improve the noise reduction effect.

[0185] The present application obtains communication audio data; obtains a first noise reduction gain for the communication audio data according to a first noise reduction model, and obtains a second noise reduction gain for the communication audio data according to a second noise reduction model; the noise reduction strength of the first noise reduction model is greater than the noise reduction strength of the second noise reduction model; the degree of voice damage to the communication audio data by the first noise reduction model is greater than the degree of voice damage to the communication audio data by the second noise reduction model; determines a combined noise reduction gain for the communication audio data according to the first noise reduction gain and the second noise reduction gain; performs noise reduction processing on the communication audio data according to the combined noise reduction gain to obtain noise-reduced audio data of the communication audio data. It can be seen from this that the method proposed in the present application can obtain a combined noise reduction gain for the communication audio data by using a first noise reduction model with relatively large noise reduction strength and a second noise reduction model with relatively good voice protection capability, and then the communication audio data can be denoised by the combined noise reduction gain, so that the voice audio in the communication audio data can be damaged to a lesser extent while the noise audio in the communication audio data is denoised to a greater extent.

[0186] See also Fig.10 , is a schematic diagram of the structure of an audio data noise reduction device provided by the present application. Fig.10 As shown, the audio data noise reduction device 1 may include: an audio acquisition module 101, a gain acquisition module 102, a gain merging module 103 and a noise reduction module 104;

[0187] The audio acquisition module 101 is used to acquire communication audio data;

[0188] A gain acquisition module 102 is used to acquire a first noise reduction gain for communication audio data according to a first noise reduction model, and to acquire a second noise reduction gain for communication audio data according to a second noise reduction model; the noise reduction strength of the first noise reduction model is greater than the noise reduction strength of the second noise reduction model; the degree of voice damage to the communication audio data by the first noise reduction model is greater than the degree of voice damage to the communication audio data by the second noise reduction model;

[0189] A gain combining module 103, configured to determine a combined noise reduction gain for the communication audio data according to the first noise reduction gain and the second noise reduction gain;

[0190] The noise reduction module 104 is used to perform noise reduction processing on the communication audio data according to the combined noise reduction gain to obtain noise-reduced audio data of the communication audio data.

[0191] The specific functional implementation of the audio acquisition module 101, the gain acquisition module 102, the gain merging module 103 and the noise reduction module 104 can be found in Figure 3 Steps S101 to S104 in the corresponding embodiment will not be described in detail here.

[0192] The gain acquisition module 102 includes: a frequency domain signal acquisition unit 1021 and a gain acquisition unit 1022;

[0193] The frequency domain signal acquisition unit 1021 is used to acquire the audio time domain signal of the communication audio data, and obtain the audio frequency domain signal of the communication audio data according to the audio time domain signal;

[0194] The gain acquisition unit 1022 is used to input the audio frequency domain signal into the first noise reduction model to obtain a first noise reduction gain, and input the audio frequency domain signal into the second noise reduction model to obtain a second noise reduction gain.

[0195] The specific functional implementation of the frequency domain signal acquisition unit 1021 and the gain acquisition unit 1022 can be found in Figure 3 The corresponding step S102 in the embodiment will not be described in detail here.

[0196] The gain merging module 103 includes: a noise estimation unit 1031 and a gain merging unit 1032;

[0197] A noise estimation unit 1031 is used to perform a noise estimation operation on the communication audio data to obtain a speech estimation probability of the communication audio data;

[0198] The gain combining unit 1032 is used to determine a combined noise reduction gain for the communication audio data according to the speech estimation probability, the first noise reduction gain and the second noise reduction gain.

[0199] The specific functional implementation of the noise estimation unit 1031 and the gain merging unit 1032 can be found in Figure 3 The corresponding step S103 in the embodiment will not be described in detail here.

[0200] The gain combining unit 1032 includes: a coefficient generating subunit 10321, a first weighting subunit 10322, a second weighting subunit 10323 and a gain determining subunit 10324;

[0201] A coefficient generating subunit 10321, used to generate a noise weighting coefficient corresponding to the speech estimation probability;

[0202] A first weighting subunit 10322 is used to weight the first noise reduction gain according to the noise weighting coefficient to obtain a noise weighted gain;

[0203] A second weighting subunit 10323, configured to weight the second noise reduction gain according to the speech estimation probability to obtain a speech weighted gain;

[0204] The gain determination subunit 10324 is used to determine the combined noise reduction gain according to the noise weighted gain and the speech weighted gain.

[0205] The specific functional implementation of the coefficient generating subunit 10321, the first weighting subunit 10322, the second weighting subunit 10323 and the gain determining subunit 10324 can be found in Figure 3 The corresponding step S103 in the embodiment will not be described in detail here.

[0206] The noise weighting coefficient includes weighting coefficients corresponding to at least two frequency points respectively; the first noise reduction gain includes noise reduction gains corresponding to at least two frequency points respectively; the noise reduction gains corresponding to at least two frequency points included in the first noise reduction gain correspond to the weighting coefficients corresponding to at least two frequency points respectively;

[0207] The first weighting subunit 10322 includes: a first frequency point weighting subunit 103221 and a first frequency point gain determining subunit 103222;

[0208] The first frequency weighting subunit 103221 is used to weight the noise reduction gains belonging to the same frequency point in the first noise reduction gains according to the weighting coefficients corresponding to at least two frequency points in the noise weighting coefficient, so as to obtain the first weighted gain corresponding to each frequency point;

[0209] The first frequency point gain determining subunit 103222 is used to determine the first weighted gain corresponding to each frequency point as the noise weighted gain.

[0210] The specific functional implementation of the first frequency point weighting subunit 103221 and the first frequency point gain determination subunit 103222 can be found in Figure 3 The corresponding step S103 in the embodiment will not be described in detail here.

[0211] The speech estimation probability includes speech probabilities corresponding to at least two frequency points respectively; the second noise reduction gain includes noise reduction gains corresponding to at least two frequency points respectively; the noise reduction gains corresponding to at least two frequency points included in the second noise reduction gain correspond to the speech probabilities corresponding to at least two frequency points respectively in a one-to-one manner;

[0212] The second weighting subunit 10323 includes: a second frequency point weighting subunit 103231 and a second frequency point gain determining subunit 103232;

[0213] The second frequency point weighting subunit 103231 is used to weight the noise reduction gains belonging to the same frequency point in the second noise reduction gains according to the weighting coefficients corresponding to the at least two frequency points in the speech estimation probability, so as to obtain the second weighted gain corresponding to each frequency point;

[0214] The second frequency point gain determining subunit 103232 is used to determine the second weighted gain corresponding to each frequency point as the speech weighted gain.

[0215] The specific functional implementation of the second frequency point weighting subunit 103231 and the second frequency point gain determination subunit 103232 can be found in Figure 3 The corresponding step S103 in the embodiment will not be described in detail here.

[0216] The combined noise reduction gain includes noise reduction gains corresponding to at least two frequency points respectively; the audio frequency domain signal of the communication audio data includes energy values ​​corresponding to at least two frequency points respectively; the noise reduction gains corresponding to at least two frequency points included in the combined noise reduction gain correspond to the energy values ​​corresponding to at least two frequency points respectively in a one-to-one manner;

[0217] The noise reduction module 104 includes: an energy value weighting unit 1041, a weighted signal determining unit 1042 and a domain transformation unit 1043;

[0218] The energy value weighting unit 1041 is used to weight the energy values ​​belonging to the same frequency point in the communication audio data according to the noise reduction gains corresponding to at least two frequency points in the combined noise reduction gain, so as to obtain the weighted energy value corresponding to each frequency point;

[0219] The weighted signal determining unit 1042 is used to determine the weighted audio frequency domain signal of the communication audio data according to the weighted energy value corresponding to each frequency point;

[0220] The domain transformation unit 1043 is used to perform time domain transformation on the weighted audio frequency domain signal to obtain noise reduction audio data of the communication audio data.

[0221] The specific functional implementation of the energy value weighting unit 1041, the weighted signal determining unit 1042 and the domain transforming unit 1043 can be found in Figure 3 The corresponding step S104 in the embodiment will not be described in detail here.

[0222] The audio data noise reduction device 1 further includes:

[0223] The first sample acquisition module 105 is used to acquire pure speech sample audio data and pure noise sample audio data; the sample audio frequency domain signal of the pure speech sample audio data includes a sample speech energy value; the sample audio frequency domain signal of the pure noise sample audio data includes a sample noise energy value;

[0224] A first actual gain acquisition module 106, configured to obtain sample actual noise reduction gains corresponding to pure speech sample audio data and pure noise sample audio data according to the sample speech energy value and the sample noise energy value;

[0225] A first predicted gain acquisition module 107 is used to synchronously input the sample actual noise reduction gain, the pure speech sample audio data and the pure noise sample audio data into a first initial noise reduction model, and predict the first sample predicted noise reduction gain corresponding to the pure speech sample audio data and the pure noise sample audio data based on the first initial noise reduction model;

[0226] The first parameter adjustment module 108 is used to adjust the model parameters of the first initial denoising model based on the actual noise reduction gain of the sample, the first sample predicted noise reduction gain and the first cost function to obtain the first noise reduction model; the first cost function is used to make the first sample predicted noise reduction gain predicted by the first initial denoising model approach the square term of the actual noise reduction gain of the sample.

[0227] The specific functional implementation of the energy value weighting unit 1041, the weighted signal determining unit 1042 and the domain transforming unit 1043 can be found in Figure 3 The corresponding step S101 in the embodiment will not be described in detail here.

[0228] The audio data denoising device further includes: a second sample acquisition module 109, a second actual gain acquisition module 110, a second predicted gain acquisition module 111, and a second parameter adjustment module 112;

[0229] The second sample acquisition module 109 is used to acquire pure speech sample audio data and pure noise sample audio data; the sample audio frequency domain signal of the pure speech sample audio data includes the sample speech energy value; the sample audio frequency domain signal of the pure noise sample audio data includes the sample noise energy value;

[0230] A second actual gain acquisition module 110 is used to obtain sample actual noise reduction gains corresponding to pure speech sample audio data and pure noise sample audio data according to the sample speech energy value and the sample noise energy value;

[0231] The second prediction gain acquisition module 111 is used to synchronously input the sample actual noise reduction gain, the pure speech sample audio data and the pure noise sample audio data into the second initial noise reduction model, and predict the second sample prediction noise reduction gain corresponding to the pure speech sample audio data and the pure noise sample audio data based on the second initial noise reduction model;

[0232] The second parameter adjustment module 112 is used to adjust the model parameters of the second initial denoising model based on the actual denoising gain of the sample, the second sample predicted denoising gain and the second cost function to obtain the second denoising model; the second cost function is used to make the second sample predicted denoising gain predicted by the second initial denoising model approach the actual denoising gain of the sample.

[0233] The specific functional implementation of the second sample acquisition module 109, the second actual gain acquisition module 110, the second predicted gain acquisition module 111 and the second parameter adjustment module 112 can be found in Figure 3 The corresponding step S101 in the embodiment will not be described in detail here.

[0234] The audio data noise reduction device 1 is also used for:

[0235] The noise reduction audio data of the communication audio data is synchronized to the connected conversation terminal so that the conversation terminal outputs the noise reduction audio data.

[0236] The present application obtains communication audio data; obtains a first noise reduction gain for the communication audio data according to a first noise reduction model, and obtains a second noise reduction gain for the communication audio data according to a second noise reduction model; the noise reduction strength of the first noise reduction model is greater than the noise reduction strength of the second noise reduction model; the degree of voice damage to the communication audio data by the first noise reduction model is greater than the degree of voice damage to the communication audio data by the second noise reduction model; determines a combined noise reduction gain for the communication audio data according to the first noise reduction gain and the second noise reduction gain; performs noise reduction processing on the communication audio data according to the combined noise reduction gain to obtain noise-reduced audio data of the communication audio data. It can be seen from this that the method proposed in the present application can obtain a combined noise reduction gain for the communication audio data by using a first noise reduction model with relatively large noise reduction strength and a second noise reduction model with relatively good voice protection capability, and then the communication audio data can be denoised by the combined noise reduction gain, so that the voice audio in the communication audio data can be damaged to a lesser extent while the noise audio in the communication audio data is denoised to a greater extent.

[0237] See also Fig.11 , is a schematic diagram of the structure of a computer device provided by this application. Fig.11As shown, the computer device 1000 may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, the computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the user interface 1003 may optionally include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 1005 may optionally be at least one storage device located away from the aforementioned processor 1001. As Fig.11 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a device control application program.

[0238] exist Fig.11 In the computer device 1000 shown in the figure, the network interface 1004 can provide a network communication function; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to implement the above Figure 3 It should be understood that the computer device 1000 described in this application can also execute the above Fig.10 The description of the audio data noise reduction device 1 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated here either.

[0239] In addition, it should be pointed out here that: the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the audio data noise reduction device 1 mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, it can execute the above-mentioned Figure 3 The description of the audio data noise reduction method in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated. For technical details not disclosed in the computer storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0240] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the above-mentioned program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the above-mentioned storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0241] The above disclosure is only the preferred embodiment of the present application, which certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A method for reducing noise in audio data, characterized in that: include: Acquiring communication audio data, the communication audio data including a target user's voice and noise other than the target user's voice; Obtaining a first noise reduction gain for the communication audio data according to a first noise reduction model, and obtaining a second noise reduction gain for the communication audio data according to a second noise reduction model; the first noise reduction model is obtained by training based on a first cost function, and the second noise reduction model is obtained by training based on a second cost function, the first cost function is used to improve the noise reduction strength of the first noise reduction model, and the second cost function is used to reduce the degree of speech damage of the second noise reduction model; The noise reduction strength of the first noise reduction model is greater than the noise reduction strength of the second noise reduction model; the degree of voice damage to the communication audio data caused by the first noise reduction model is greater than the degree of voice damage to the communication audio data caused by the second noise reduction model; Performing a noise estimation operation on the communication audio data to obtain a speech estimation probability of the communication audio data; Generating a noise weighting coefficient corresponding to the speech estimation probability; Weighting the first noise reduction gain according to the noise weighting coefficient to obtain a noise weighted gain; Weighting the second noise reduction gain according to the speech estimation probability to obtain a speech weighted gain; Determining a combined noise reduction gain for the communication audio data according to the noise weighted gain and the speech weighted gain; The communication audio data is subjected to noise reduction processing according to the combined noise reduction gain to obtain noise-reduced audio data of the communication audio data.

2. The method according to claim 1, characterized in that The step of acquiring a first noise reduction gain for the communication audio data according to the first noise reduction model, and acquiring a second noise reduction gain for the communication audio data according to the second noise reduction model comprises: Acquire an audio time domain signal of the communication audio data, and obtain an audio frequency domain signal of the communication audio data according to the audio time domain signal; The audio frequency domain signal is input into the first noise reduction model to obtain the first noise reduction gain, and the audio frequency domain signal is input into the second noise reduction model to obtain the second noise reduction gain.

3. The method according to claim 1, characterized in that The noise weighting coefficient includes weighting coefficients corresponding to at least two frequency points respectively; the first noise reduction gain includes noise reduction gains corresponding to the at least two frequency points respectively; the noise reduction gains corresponding to the at least two frequency points included in the first noise reduction gain correspond to the weighting coefficients corresponding to the at least two frequency points respectively in a one-to-one correspondence; The step of weighting the first noise reduction gain according to the noise weighting coefficient to obtain a noise weighted gain includes: According to the weighting coefficients corresponding to the at least two frequency points in the noise weighting coefficient, respectively weighting the noise reduction gains belonging to the same frequency point in the first noise reduction gains, to obtain first weighted gains corresponding to each frequency point; The first weighted gain corresponding to each frequency point is determined as the noise weighted gain.

4. The method according to claim 1, characterized in that: The speech estimation probability includes speech probabilities corresponding to at least two frequency points respectively; the second noise reduction gain includes noise reduction gains corresponding to the at least two frequency points respectively; the noise reduction gains corresponding to the at least two frequency points respectively included in the second noise reduction gain correspond to the speech probabilities corresponding to the at least two frequency points respectively in a one-to-one manner; The step of weighting the second noise reduction gain according to the speech estimation probability to obtain a speech weighted gain includes: According to the weighting coefficients respectively corresponding to the at least two frequency points in the speech estimation probability, the noise reduction gains belonging to the same frequency point in the second noise reduction gains are weighted respectively to obtain second weighted gains respectively corresponding to each frequency point; The second weighted gain corresponding to each frequency point is determined as the speech weighted gain.

5. The method according to claim 1, characterized in that The combined noise reduction gain includes noise reduction gains corresponding to at least two frequency points respectively; the audio frequency domain signal of the communication audio data includes energy values ​​corresponding to the at least two frequency points respectively; the noise reduction gains corresponding to the at least two frequency points respectively included in the combined noise reduction gain correspond to the energy values ​​corresponding to the at least two frequency points respectively in a one-to-one manner; The performing noise reduction processing on the communication audio data according to the combined noise reduction gain to obtain noise-reduced audio data of the communication audio data includes: According to the noise reduction gains corresponding to the at least two frequency points in the combined noise reduction gain, respectively, weighting the energy values ​​belonging to the same frequency point in the communication audio data to obtain a weighted energy value corresponding to each frequency point; Determining a weighted audio frequency domain signal of the communication audio data according to the weighted energy value corresponding to each frequency point; The weighted audio frequency domain signal is transformed in the time domain to obtain the noise reduction audio data of the communication audio data.

6. The method according to claim 1, characterized in that Also includes: Acquire pure speech sample audio data and pure noise sample audio data; the sample audio frequency domain signal of the pure speech sample audio data includes a sample speech energy value; the sample audio frequency domain signal of the pure noise sample audio data includes a sample noise energy value; Obtaining actual noise reduction gains of samples corresponding to the pure speech sample audio data and the pure noise sample audio data according to the sample speech energy value and the sample noise energy value; The sample actual noise reduction gain, the pure speech sample audio data and the pure noise sample audio data are synchronously input into the first initial noise reduction model, and based on the first initial noise reduction model, a first sample predicted noise reduction gain corresponding to the pure speech sample audio data and the pure noise sample audio data is predicted; Based on the actual noise reduction gain of the sample, the predicted noise reduction gain of the first sample and a first cost function, adjusting the model parameters of the first initial noise reduction model to obtain the first noise reduction model; The first cost function is used to make the first sample predicted noise reduction gain predicted by the first initial noise reduction model approach the square term of the sample actual noise reduction gain.

7. The method according to claim 6, characterized in that Also includes: Acquire pure speech sample audio data and pure noise sample audio data; the sample audio frequency domain signal of the pure speech sample audio data includes a sample speech energy value; the sample audio frequency domain signal of the pure noise sample audio data includes a sample noise energy value; Obtaining actual noise reduction gains of samples corresponding to the pure speech sample audio data and the pure noise sample audio data according to the sample speech energy value and the sample noise energy value; The sample actual noise reduction gain, the pure speech sample audio data and the pure noise sample audio data are synchronously input into the second initial noise reduction model, and the second sample predicted noise reduction gain corresponding to the pure speech sample audio data and the pure noise sample audio data is predicted based on the second initial noise reduction model; Based on the actual noise reduction gain of the sample, the predicted noise reduction gain of the second sample and the second cost function, adjusting the model parameters of the second initial noise reduction model to obtain the second noise reduction model; The second cost function is used to make the second sample predicted noise reduction gain predicted by the second initial noise reduction model approach the sample actual noise reduction gain.

8. The method according to claim 1, characterized in that Also includes: The noise reduction audio data of the communication audio data is synchronized to a connected conversation terminal so that the conversation terminal outputs the noise reduction audio data.

9. An audio data noise reduction device, characterized in that: include: An audio acquisition module, used to acquire communication audio data, wherein the communication audio data includes a target user's voice and noise other than the target user's voice; A gain acquisition module, used to acquire a first noise reduction gain for the communication audio data according to a first noise reduction model, and to acquire a second noise reduction gain for the communication audio data according to a second noise reduction model; the first noise reduction model is obtained by training based on a first cost function, and the second noise reduction model is obtained by training based on a second cost function, the first cost function is used to improve the noise reduction strength of the first noise reduction model, and the second cost function is used to reduce the degree of speech damage of the second noise reduction model; The noise reduction strength of the first noise reduction model is greater than the noise reduction strength of the second noise reduction model; the degree of voice damage to the communication audio data caused by the first noise reduction model is greater than the degree of voice damage to the communication audio data caused by the second noise reduction model; a gain combining module, configured to determine a combined noise reduction gain for the communication audio data according to the first noise reduction gain and the second noise reduction gain; A noise reduction module, configured to perform noise reduction processing on the communication audio data according to the combined noise reduction gain to obtain noise-reduced audio data of the communication audio data; The gain combining module comprises: A noise estimation unit, configured to perform a noise estimation operation on the communication audio data to obtain a speech estimation probability of the communication audio data; a gain combining unit, configured to determine the combined noise reduction gain for the communication audio data according to the speech estimation probability, the first noise reduction gain, and the second noise reduction gain; The gain combining unit comprises: A coefficient generating subunit, used for generating a noise weighting coefficient corresponding to the speech estimation probability; A first weighting subunit, configured to weight the first noise reduction gain according to the noise weighting coefficient to obtain a noise weighted gain; A second weighting subunit, configured to weight the second noise reduction gain according to the speech estimation probability to obtain a speech weighted gain; The gain determination subunit is used to determine the combined noise reduction gain according to the noise weighted gain and the speech weighted gain.

10. The device according to claim 9, characterized in that The gain acquisition module comprises: A frequency domain signal acquisition unit, configured to acquire an audio time domain signal of the communication audio data, and obtain an audio frequency domain signal of the communication audio data according to the audio time domain signal; The gain acquisition unit is used to input the audio frequency domain signal into the first noise reduction model to obtain the first noise reduction gain, and input the audio frequency domain signal into the second noise reduction model to obtain the second noise reduction gain.

11. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the method according to any one of claims 1 to 8 is executed.

Citation Information

Patent Citations

  • Apparatus and method for providing informed multichannel speech presence probability estimation

    CN104781880A

  • Microphone expansion for background noise reduction

    CN1164171A

  • Gain scaling for higher signal-to-noise ratios in multistage, multi-bit delta sigma modulators

    US20020093442A1