Audio noise reduction, audio noise reduction model processing method, device, equipment and medium
By introducing a frequency domain processing sub-model to the audio noise reduction model for frequency domain feature encoding, and combining the time domain processing sub-model for noise reduction, the problem of failure to fully utilize frequency domain information in the prior art is solved, and a better audio noise reduction effect is achieved.
Patent Information
- Application Number
- CN202110557785.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-05-21
AI Technical Summary
The existing audio noise reduction model fails to fully utilize the frequency domain information of the audio signal, resulting in poor noise reduction.
By training the real-part processing network and the imaginary processing network in the frequency domain processing sub-model, the original audio signal is subject to frequency domain transformation and feature encoding, real-part attention and imaginary attention, combined with the real-part sequence and imaginary sequence, frequency-domain encoding features are generated, and noise reduction is performed through the time domain processing sub-model.
The audio noise reduction effect is improved. By fully utilizing the frequency domain information of the audio signal, the characterization of clean audio signals is enhanced, and the sound quality of the time domain signals is further improved.
Smart Images

Figure CN113763979B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an audio noise reduction method, apparatus, computer equipment and storage medium, and also to a processing method, apparatus, computer equipment and storage medium for an audio noise reduction model. Background Art
[0002] As we all know, audio signals are generally mixed with noise to varying degrees. For example, when a user is making an audio call, they may be in various scenarios, and noisy background noise will interfere with the audio call. In order to obtain better audio quality, the original audio signal is usually subjected to noise reduction. Traditional audio noise reduction methods include adaptive filters, spectral subtraction, and Wiener filtering.
[0003] With the popularity of deep learning in the field of artificial intelligence technology, the use of deep learning models based on neural networks to reduce the noise of audio signals has become a research hotspot, and its effect is better than traditional noise reduction algorithms. However, some existing audio noise reduction models do not maximize the use of the frequency domain information of the input audio signal, resulting in poor audio noise reduction effects. Summary of the invention
[0004] Based on this, it is necessary to provide an audio noise reduction method, device, computer equipment and storage medium that can improve the audio noise reduction effect in response to the above-mentioned technical problems, and also provide a processing method, device, computer equipment and storage medium for an audio noise reduction model that can improve the audio noise reduction effect.
[0005] An audio noise reduction method, the method comprising:
[0006] Get the original audio signal in the time domain;
[0007] The real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model are used to respectively perform feature encoding on the real part sequence and the imaginary part sequence obtained after the original audio signal is transformed into the frequency domain signal, so as to obtain the real part attention and the imaginary part attention corresponding to the original audio signal;
[0008] Based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, obtaining a frequency domain coding feature corresponding to the original audio signal;
[0009] The frequency domain coding features are transformed into time domain signals, and the time domain signals are subjected to noise reduction processing through a trained time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal.
[0010] An audio noise reduction device, comprising:
[0011] An acquisition module, used for acquiring the original audio signal in the time domain;
[0012] A frequency domain coding module is used to perform feature coding on the real part sequence and the imaginary part sequence obtained after the original audio signal is transformed into the frequency domain signal through the real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model, respectively, to obtain the real part attention and the imaginary part attention corresponding to the original audio signal; based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, obtain the frequency domain coding feature corresponding to the original audio signal;
[0013] The time domain noise reduction module is used to transform the frequency domain coding features into a time domain signal, and perform noise reduction processing on the time domain signal through a trained time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal.
[0014] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0015] Get the original audio signal in the time domain;
[0016] The real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model are used to respectively perform feature encoding on the real part sequence and the imaginary part sequence obtained after the original audio signal is transformed into the frequency domain signal, so as to obtain the real part attention and the imaginary part attention corresponding to the original audio signal;
[0017] Based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, obtaining a frequency domain coding feature corresponding to the original audio signal;
[0018] The frequency domain coding features are transformed into time domain signals, and the time domain signals are subjected to noise reduction processing through a trained time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal.
[0019] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0020] Get the original audio signal in the time domain;
[0021] The real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model are used to respectively perform feature encoding on the real part sequence and the imaginary part sequence obtained after the original audio signal is transformed into the frequency domain signal, so as to obtain the real part attention and the imaginary part attention corresponding to the original audio signal;
[0022] Based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, obtaining a frequency domain coding feature corresponding to the original audio signal;
[0023] The frequency domain coding features are transformed into time domain signals, and the time domain signals are subjected to noise reduction processing through a trained time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal.
[0024] A computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps of the above-mentioned audio noise reduction method.
[0025] In the above-mentioned audio noise reduction method, device, computer equipment and storage medium, the trained frequency domain processing sub-model includes a real part processing network and an imaginary part processing network. After obtaining the original audio signal in the time domain, the real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model are used to respectively perform feature encoding on the real part sequence and the imaginary part sequence of the original audio signal. Such a network structure can make full use of the frequency domain information of the original audio signal, that is, the amplitude information and phase information represented by the real part sequence and the imaginary part sequence, so that the real part attention and the imaginary part attention obtained by encoding can be used for the original audio signal. More and more accurate attention is given to the clean audio signal, so that the frequency domain coding features corresponding to the original audio signal obtained based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention can accurately characterize the frequency domain features of the clean audio signal in the original audio signal, and the time domain signal obtained according to the frequency domain coding features can also accurately express the clean audio signal in the original audio signal, and the noise reduction effect is better; in addition, by further using the time domain processing sub-model to perform noise reduction on the time domain signal, the sound quality of the time domain signal can be further improved, and the effect of the noise reduction signal obtained will also be better.
[0026] A method for processing an audio noise reduction model, the method comprising:
[0027] Acquire a sample audio signal, wherein the sample audio information is generated according to a clean audio signal;
[0028] By using a real part processing network and an imaginary part processing network in a first sub-model based on a neural network, feature encoding is performed on a real part sequence and an imaginary part sequence obtained after the sample audio signal is transformed into a frequency domain signal, thereby obtaining a real part attention and an imaginary part attention corresponding to the sample audio signal, and based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, a frequency domain coding feature corresponding to the sample audio signal is obtained;
[0029] Performing model training on the first sub-model according to a first loss determined based on the frequency domain coding feature and a frequency domain transformation sequence corresponding to the clean audio signal to obtain a frequency domain processing sub-model;
[0030] The frequency domain processing sub-model is connected to the time domain processing sub-model to be trained and trained together to obtain the audio noise reduction model for performing noise reduction processing on the audio signal.
[0031] In one embodiment, the real part processing network and the imaginary part processing network in the first sub-model based on the neural network respectively perform feature encoding on the real part sequence and the imaginary part sequence obtained after the sample audio signal is transformed into a frequency domain signal to obtain the real part attention and the imaginary part attention corresponding to the sample audio signal, including:
[0032] Inputting the sample audio signal into a first sub-model based on a neural network;
[0033] In the first sub-model, frequency domain transformation is performed on the sample audio signal to obtain a real part sequence and an imaginary part sequence corresponding to the sample audio signal;
[0034] Through the real part processing network in the first sub-model, feature encoding is performed on the real part sequence and the imaginary part sequence respectively to obtain a real part first encoding feature and an imaginary part first encoding feature;
[0035] Through the imaginary part processing network in the first sub-model, feature encoding is performed on the real part sequence and the imaginary part sequence respectively to obtain a real part second encoding feature and an imaginary part second encoding feature;
[0036] According to the real part first coding feature and the imaginary part second coding feature, the real part attention corresponding to the sample audio signal is obtained, and according to the real part second coding feature and the imaginary part first coding feature, the imaginary part attention corresponding to the sample audio signal is obtained.
[0037] In one embodiment, obtaining the frequency domain coding feature corresponding to the sample audio signal based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention includes:
[0038] Multiplying the real part sequence by the real part attention to obtain the real part of the frequency domain coding feature corresponding to the sample audio signal;
[0039] The imaginary part sequence is multiplied by the imaginary part attention to obtain the imaginary part of the frequency domain coding feature corresponding to the sample audio signal.
[0040] In one embodiment, obtaining the frequency domain coding feature corresponding to the sample audio signal based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention includes:
[0041] Multiplying the real part sequence by the real part attention to obtain a first result, multiplying the imaginary part sequence by the imaginary part attention to obtain a second result, and taking the difference between the first result and the second result as the real part of the frequency domain coding feature corresponding to the sample audio signal;
[0042] The real part sequence is multiplied by the imaginary part attention to obtain a third result, the imaginary part sequence is multiplied by the real part attention to obtain a fourth result, and the sum of the third result and the fourth result is used as the imaginary part of the frequency domain coding feature corresponding to the sample audio signal.
[0043] In one embodiment, obtaining the frequency domain coding feature corresponding to the sample audio signal based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention includes:
[0044] Based on the real part sequence and the imaginary part sequence, original amplitude information and original phase information of the sample audio signal are obtained; based on the real part attention and the imaginary part attention, predicted amplitude information and predicted phase information of the sample audio signal are obtained;
[0045] Obtaining amplitude information of a frequency domain coding feature corresponding to the sample audio signal according to the product of the original amplitude information and the predicted amplitude information;
[0046] The phase information of the frequency domain coding feature corresponding to the sample audio signal is obtained according to the sum of the original phase information and the predicted phase information.
[0047] In one embodiment, the step of determining the first loss includes:
[0048] Performing frequency domain transformation processing on the clean audio signal to obtain a corresponding frequency domain transformation sequence, wherein the frequency domain transformation sequence includes a real part sequence and an imaginary part sequence;
[0049] The first loss is determined based on the difference between the real part sequence corresponding to the clean audio signal and the real part features in the frequency domain coding features corresponding to the sample audio signal, and the difference between the imaginary part sequence corresponding to the clean audio signal and the imaginary part features in the frequency domain coding features corresponding to the sample audio signal.
[0050] In one embodiment, the step of determining the second loss comprises:
[0051] The encoder in the second sub-model encodes the time domain signal to obtain a time domain code vector, the temporal feature extraction network in the second sub-model extracts features from the time domain code vector to obtain hidden features corresponding to the time domain signal, and the decoder in the second sub-model decodes based on the time domain code vector and the hidden features to obtain a noise reduction signal corresponding to the sample audio signal in the second sample set;
[0052] A second loss is constructed based on the noise reduction signal and the clean audio signal.
[0053] In one embodiment, constructing a second loss according to the noise reduction signal and the clean audio signal comprises:
[0054] Projecting the noise reduction signal corresponding to the sample audio signal in the vertical direction and the horizontal direction of the clean audio signal respectively to obtain a vertical projection vector and a horizontal projection vector;
[0055] A second loss is obtained according to the vertical projection vector and the horizontal projection vector.
[0056] In one embodiment, the step of determining the third loss includes:
[0057] The encoder in the third sub-model is used to encode the time domain signal to obtain a time domain code vector, the time series feature extraction network in the third sub-model is used to extract features from the time domain code vector to obtain hidden features corresponding to the time domain signal, and the output layer in the third sub-model is used to predict the noise scene category of the sample audio signal based on the hidden features;
[0058] A third loss is constructed according to the noise scene category and the noise label category of the noise signal used to generate the sample audio signal.
[0059] A processing device for an audio noise reduction model, the device comprising:
[0060] An acquisition module, used for acquiring a sample audio signal, wherein the sample audio information is generated according to a clean audio signal;
[0061] A frequency domain coding training module is used to perform feature coding on a real part sequence and an imaginary part sequence obtained after transforming the sample audio signal into a frequency domain signal, respectively, through a real part processing network and an imaginary part processing network in a first sub-model based on a neural network, to obtain a real part attention and an imaginary part attention corresponding to the sample audio signal, and to obtain a frequency domain coding feature corresponding to the sample audio signal based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention; and to perform model training on the first sub-model according to a first loss determined based on the frequency domain coding feature and a frequency domain transformation sequence corresponding to the clean audio signal, to obtain a frequency domain processing sub-model;
[0062] The integrated training module is used to connect the frequency domain processing sub-model and the time domain processing sub-model to be trained and train them together to obtain the audio noise reduction model used for noise reduction processing of audio signals.
[0063] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0064] Acquire a sample audio signal, wherein the sample audio information is generated according to a clean audio signal;
[0065] By using a real part processing network and an imaginary part processing network in a first sub-model based on a neural network, feature encoding is performed on a real part sequence and an imaginary part sequence obtained after the sample audio signal is transformed into a frequency domain signal, thereby obtaining a real part attention and an imaginary part attention corresponding to the sample audio signal, and based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, a frequency domain coding feature corresponding to the sample audio signal is obtained;
[0066] Performing model training on the first sub-model according to a first loss determined based on the frequency domain coding feature and a frequency domain transformation sequence corresponding to the clean audio signal to obtain a frequency domain processing sub-model;
[0067] The frequency domain processing sub-model is connected to the time domain processing sub-model to be trained and trained together to obtain the audio noise reduction model for performing noise reduction processing on the audio signal.
[0068] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0069] Acquire a sample audio signal, wherein the sample audio information is generated according to a clean audio signal;
[0070] By using a real part processing network and an imaginary part processing network in a first sub-model based on a neural network, feature encoding is performed on a real part sequence and an imaginary part sequence obtained after the sample audio signal is transformed into a frequency domain signal, thereby obtaining a real part attention and an imaginary part attention corresponding to the sample audio signal, and based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, a frequency domain coding feature corresponding to the sample audio signal is obtained;
[0071] Performing model training on the first sub-model according to a first loss determined based on the frequency domain coding feature and a frequency domain transformation sequence corresponding to the clean audio signal to obtain a frequency domain processing sub-model;
[0072] The frequency domain processing sub-model is connected to the time domain processing sub-model to be trained and trained together to obtain the audio noise reduction model for performing noise reduction processing on the audio signal.
[0073] A computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the above-mentioned audio noise reduction model processing method.
[0074] The processing method, device, computer equipment and storage medium of the above-mentioned audio noise reduction model, the audio noise reduction model includes a frequency domain processing sub-model and a time domain processing sub-model, wherein the frequency domain processing sub-model includes a real part processing network and an imaginary part processing network. After obtaining the original audio signal in the time domain, the real part processing network and the imaginary part processing network in the frequency domain processing sub-model are used to perform feature encoding on the real part sequence and the imaginary part sequence of the sample audio signal respectively. Such a network structure can fully learn the frequency domain information of the original audio signal, that is, the amplitude information and phase information represented by the real part sequence and the imaginary part sequence, so that the real part attention and the imaginary part attention obtained by encoding can be It is able to give more and more accurate attention to the clean audio signal in the sample audio signal. In this way, based on the frequency domain coding features corresponding to the sample audio signal obtained by the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, the frequency domain processing sub-model is trained according to the first loss determined by the frequency domain coding features and the frequency domain transformation sequence corresponding to the clean audio signal that generates the sample audio signal. The frequency domain processing sub-model can accurately learn the frequency domain features of the clean signal in the sample audio signal. Subsequently, the frequency domain processing sub-model is combined with the time domain processing sub-model for model training, and the denoising effect of the obtained audio denoising model will be better. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 A diagram showing an application environment of an audio noise reduction method in an embodiment;
[0076] Figure 2 A schematic diagram of a process of a method for processing an audio noise reduction model in one embodiment;
[0077] Figure 3 A schematic diagram of a model framework of a frequency domain processing sub-model in one embodiment;
[0078] Figure 4 A schematic diagram of cascading a frequency domain processing sub-model and a time domain processing sub-model to obtain an audio noise reduction model in one embodiment;
[0079] Figure 5 A schematic diagram of a flow chart of obtaining the real part attention and the imaginary part attention corresponding to the sample audio signal in one embodiment;
[0080] Figure 6 A schematic diagram of a real part processing network and an imaginary part processing network performing feature encoding on a real part sequence and an imaginary part sequence in one embodiment;
[0081] Figure 7 A schematic diagram of a two-layer real part processing network and an imaginary part processing network in one embodiment;
[0082] Figure 8 A schematic diagram of a process of connecting a frequency domain processing sub-model and a time domain processing sub-model to be trained and jointly training them to obtain an audio noise reduction model in an embodiment;
[0083] FIG9( a ) is a schematic diagram of a model structure of a second sub-model in one embodiment;
[0084] FIG9( b ) is a schematic diagram showing the relationship between the second loss and the projection vector in one embodiment;
[0085] Fig.10 is a schematic diagram of the model structure of the third sub-model in one embodiment;
[0086] Fig.11 A schematic diagram of a model structure for performing multi-task learning on a time domain noise reduction task and a noise scene classification task in one embodiment;
[0087] Fig.12 A schematic diagram of a process of connecting a frequency domain processing sub-model and a time domain processing sub-model to be trained and training them together to obtain an audio noise reduction model in another embodiment;
[0088] Fig.13 A schematic diagram of the steps of training an audio noise reduction model in one embodiment;
[0089] Fig.14 A schematic diagram of a process of performing model distillation training on a time-domain denoising sub-model in one embodiment;
[0090] Fig.15A schematic diagram of a framework for performing model distillation training on a time-domain denoising sub-model in one embodiment;
[0091] Fig.16 is a schematic flow chart of an audio noise reduction method in one embodiment;
[0092] Fig.17 A schematic diagram of a process of obtaining the real part attention and the imaginary part attention corresponding to the original audio signal in one embodiment;
[0093] Fig.18 A schematic diagram of a process of obtaining the real part attention and the imaginary part attention corresponding to the original audio signal through the real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model in one embodiment;
[0094] Fig.19 is a structural block diagram of a processing device for an audio noise reduction model in one embodiment;
[0095] Fig. 20 is a structural block diagram of an audio noise reduction device in one embodiment;
[0096] Fig.21 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0097] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0098] The audio noise reduction method and the processing method of the audio noise reduction model provided in the present application realize audio noise reduction by using the machine learning and other technologies in the artificial intelligence (AI) technology.
[0099] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0100] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0101] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, etc. I believe that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0102] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. This discipline specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Artificial neural networks are an important machine learning technology that has broad application prospects in system identification, pattern recognition, intelligent control and other fields.
[0103] Deep learning is a new research direction in the field of artificial intelligence technology. Deep learning is the inherent laws and representation levels of learning sample data. The information obtained in the learning process is of great help in the interpretation of data such as text, images and sounds. Its ultimate goal is to enable machines to have analytical learning capabilities like humans and to recognize data such as text, images and sounds. The motivation for studying deep learning is to establish a neural network that simulates the human brain for analytical learning. It imitates the mechanism of the human brain to interpret data such as images, sounds and text. It can be understood that this application trains and uses an audio noise reduction model by using deep learning technology.
[0104] The audio noise reduction method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The terminal 102 can obtain the original audio signal in the time domain; through the real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model, the real part sequence and the imaginary part sequence obtained after the original audio signal is transformed into the frequency domain signal are feature encoded respectively, and the real part attention and the imaginary part attention corresponding to the original audio signal are obtained; based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, the frequency domain coding feature corresponding to the original audio signal is obtained; the frequency domain coding feature is transformed into a time domain signal, and the time domain signal is subjected to denoising processing through the trained time domain processing sub-model to obtain a denoised signal corresponding to the original audio signal.
[0105] For example, during an audio call, the terminal 102 can use the voice call signal as the original audio signal, and use the audio noise reduction method provided in the embodiment of the present application to perform voice noise reduction processing on the voice call signal to obtain a corresponding noise reduction signal.
[0106] It can be understood that in other embodiments, after the terminal 102 obtains the original audio signal, the original audio signal can be transmitted to the server 104. After the server 104 obtains the original audio signal, the original audio signal is input into the trained frequency domain processing sub-model, and then, after the frequency domain coding features corresponding to the original audio signal are obtained through the frequency domain processing sub-model, the frequency domain coding features are transformed into a time domain signal, and then the time domain signal is input into the trained time domain processing sub-model, and then, the time domain signal is subjected to noise reduction processing through the time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal.
[0107] The processing method of the audio noise reduction model provided in the embodiment of the present application can also be applied to Figure 1 In the application environment shown. The server 104 can obtain a sample audio signal, and the sample audio information is generated according to the clean audio signal; through the real part processing network and the imaginary part processing network in the first sub-model based on the neural network, the real part sequence and the imaginary part sequence obtained after the sample audio signal is transformed into the frequency domain signal are respectively feature encoded to obtain the real part attention and the imaginary part attention corresponding to the sample audio signal, and the frequency domain coding feature corresponding to the sample audio signal is obtained based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention; the first sub-model is model trained according to the first loss determined based on the frequency domain coding feature and the frequency domain transformation sequence corresponding to the clean audio signal to obtain the frequency domain processing sub-model; the frequency domain processing sub-model is connected with the time domain processing sub-model to be trained and trained together to obtain an audio noise reduction model for noise reduction processing of the audio signal.
[0108] It is understandable that in other embodiments, the audio noise reduction model may also be obtained by training by the terminal 102 .
[0109] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices, and the server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this.
[0110] The audio noise reduction method provided in the embodiment of the present application may be executed by the audio noise reduction device provided in the embodiment of the present application, or a computer device integrated with the audio noise reduction device, wherein the audio noise reduction device may be implemented in hardware or software. The processing method of the audio noise reduction model provided in the embodiment of the present application may be executed by the processing device of the audio noise reduction model provided in the embodiment of the present application, or a computer device integrated with the processing device of the audio noise reduction model, wherein the processing device of the audio noise reduction model may be implemented in hardware or software. The computer device may be Figure 1 The terminal 102 or server 104 shown in .
[0111] In one embodiment, the terminal 102 or server 104 used to train the speech noise reduction model can be a node in the blockchain network. Blockchain is a new application mode of computer technologies such as distributed data storage, point-to-point transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.
[0112] The underlying blockchain platform can include processing modules such as user management, basic services, smart contracts, and operation monitoring. Among them, the user management module is responsible for the identity information management of all blockchain participants, including maintaining public and private key generation (account management), key management, and the maintenance of the correspondence between the user's real identity and the blockchain address (authority management), etc., and, under authorization, supervises and audits the transactions of certain real identities and provides risk control rule configuration (risk control audit); the basic service module is deployed on all blockchain node devices to verify the validity of business requests, and records valid requests to storage after consensus is reached. For a new business request, the basic service first adapts the interface for parsing and authentication (interface adaptation), and then encrypts the business information through the consensus algorithm (consensus management). The smart contract module is responsible for the registration and issuance of contracts, as well as contract triggering and contract execution. Developers can define the contract logic in a programming language and publish it to the blockchain (contract registration). According to the logic of the contract terms, the key or other events are called to trigger the execution and complete the contract logic. It also provides the function of contract upgrade and cancellation. The operation monitoring module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation and real-time status visualization output of the product during the product release process, such as alarm, network status monitoring, node equipment health status monitoring, etc.
[0113] The platform product service layer provides the basic capabilities and implementation framework of typical applications. Developers can superimpose business features based on these basic capabilities to complete the blockchain implementation of business logic. The application service layer provides application services based on blockchain solutions for business participants to use.
[0114] In one embodiment, Figure 2 As shown, a processing method for an audio noise reduction model is provided, and the method is applied to Figure 1 The computer device (terminal 102 or server 104) in the example is used for explanation, and the following steps are included:
[0115] Step 202: Acquire a sample audio signal, where the sample audio information is generated based on the clean audio signal.
[0116] The sample audio signal is an audio signal used to train an audio noise reduction model. After a model training requirement for audio noise reduction is generated, a sample audio signal used to train the audio noise reduction model needs to be generated first.
[0117] In one embodiment, the computer device can mix the clean audio signal and the noise signal according to different signal-to-noise ratios to obtain a sample audio signal. When performing model training, the clean audio signal can be used as the label information of the generated sample audio signal. The clean audio signal can be, for example, a clean human voice signal, and the clean human voice languages used include English, Chinese, and various local dialects. The noise signal used can be noise from various different scenes, such as white noise, wind, subway sound, keyboard sound, mouse sound, etc. In some embodiments, the category of the noise signal can also be used as the label information of the sample audio signal.
[0118] Optionally, the computer device can read in the clean audio signal and the noise signal, and then randomly mix them according to different signal-to-noise ratios to obtain a sample audio signal. This can enhance the training sample data to a certain extent, thereby improving the generalization ability of the model.
[0119] In some embodiments, the computer device may also obtain sample audio signals transmitted by other computer devices, such as the above Figure 1 The server 104 obtains the image transmitted by the terminal 102, and the computer device can also obtain the sample audio signal generated on the local device.
[0120] Step 204, feature encoding is performed on the real part sequence and the imaginary part sequence obtained after converting the sample audio signal into a frequency domain signal through the real part processing network and the imaginary part processing network in the first sub-model based on the neural network, respectively, to obtain the real part attention and the imaginary part attention corresponding to the sample audio signal, and based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, the frequency domain coding feature corresponding to the sample audio signal is obtained.
[0121] Among them, the first sub-model based on the neural network can learn the characteristics of the clean audio signal in the sample audio signal in the frequency domain during the model training process through the sample audio signal. The first sub-model can adopt a deep learning model based on a neural network, such as LSTM (Long short-term memory), which is a recurrent neural network with a special structure that can learn the long-term dependencies of long sequence inputs.
[0122] Both time domain and frequency domain are basic properties of signals. Analyzing signals from different dimensions and solving problems from different angles can be called domains. The time domain reflects the correspondence between mathematical functions or physical signals and time. It is the feedback of the real world and the only objectively existing domain. The frequency domain is a coordinate system used to describe the characteristics of a signal in the frequency domain. It shows the amount of signal within a frequency range. It is a way of auxiliary thinking constructed from a mathematical perspective. Frequency domain signals can be expressed in complex form, including real and imaginary parts. Frequency domain signals can also be expressed in the form of amplitude and phase.
[0123] In order to fully learn the frequency domain characteristics of the clean audio signal mixed in the sample audio signal, the first sub-model based on the neural network includes a real part processing network and an imaginary part processing network. The real part processing network is designed to utilize the amplitude information of the sample audio signal, and the imaginary part sequence network is designed to utilize the phase information of the sample audio signal. The real part processing network and the imaginary part processing network can both be based on the LSTM (Long Short-Term Memory) network structure, and the real part processing network and the imaginary part processing network can both include at least one layer of LSTM network structure, that is, the overall LSTM structure capable of processing complex frequency domain signals includes an LSTM for processing real part sequences and an LSTM for processing imaginary part sequences, which can be called complex-LSTM.
[0124] The real part processing network can be used to obtain the real part attention corresponding to the sample audio signal, and the imaginary part processing network can be used to obtain the imaginary part attention corresponding to the sample audio signal. The real part attention can be used to reflect the attention to the clean signal in the real part frequency domain characteristics of the sample audio signal, and the imaginary part attention can be used to reflect the attention to the clean signal in the imaginary part frequency domain characteristics of the sample audio signal.
[0125] Since the goal defined in the training of the first sub-model is to multiply the real features in the output results of the real part processing network and the imaginary part processing network with the real part sequence of the original audio signal, the real part of the clean audio signal can be obtained, and the imaginary features can be multiplied with the imaginary part sequence of the original audio signal to obtain the imaginary part of the clean audio signal. Based on such a structure equivalent to the attention mechanism (Attention), the output results of the real part processing network and the imaginary part processing network can express more and more accurate attention to the clean audio signal in the sample audio signal. Therefore, the output results of the real part processing network and the imaginary part processing network can be called real attention and imaginary attention. The attention mechanism comes from the process of imitating biological observation of things. It is a mechanism that aligns internal experience with external stimuli to increase attention to some areas.
[0126] Usually, after the input signal is input into the neural network, the network parameters in the network layer of the neural network will operate on the input signal to obtain the operation result. Each layer of the network will receive the operation result output by the previous layer of the network, and then after the operation of the current layer of the network, output the operation result of the current layer as the input of the next layer. In this embodiment, the real part processing network and the imaginary part processing network in the first sub-model include at least one layer.
[0127] Specifically, the computer device can input the acquired sample audio signal into the first sub-model. In the first sub-model, the sample audio signal is first converted into a frequency domain signal, and then the real sequence and imaginary sequence of the frequency domain signal are respectively input into the real processing network and the imaginary processing network. The real processing network and the imaginary processing network will each perform feature encoding on the real sequence and the imaginary sequence, that is, the network parameters in the real processing network and the imaginary processing network operate on the input real sequence and imaginary sequence, and the real attention and imaginary attention corresponding to the sample audio signal are obtained according to the output results of the last layer of the real processing network and the imaginary processing network.
[0128] After obtaining the attention to the clean audio signal in the sample audio signal, the computer device then obtains the frequency domain coding features corresponding to the sample audio signal based on the real part sequence and imaginary part sequence of the sample audio signal, the real part attention and the imaginary part attention. The frequency domain coding features are features that reflect the audio signal from the frequency domain perspective. As mentioned above, the real part attention and the imaginary part attention output by the real part processing network and the real part processing network in the first sub-model can give the clean audio signal in the sample audio signal more and more accurate attention. Then, the computer device can dig out more parts related to the clean audio signal from the real part sequence and the imaginary part sequence of the frequency domain signal of the sample audio signal based on the real part attention and the imaginary part attention, and then obtain the frequency domain coding features corresponding to the sample audio signal. It can be understood that the frequency domain coding features include real and imaginary parts.
[0129] In one embodiment, the computer device can transform the original audio signal into a frequency domain signal using a Fourier transform in the first sub-model, and the frequency domain signal includes a real sequence and an imaginary sequence. The Fourier transform can be a short-time Fourier transform (STFT). For example, the sample audio signal is a 15s audio signal. When the sampling rate is 16K, the Fourier transform window length is set to 512, and the Fourier transform window overlap is set to 75%, that is, when the window displacement step is 128, 1872 discrete sequences with a length of 512 can be obtained according to the sample audio signal. After the short-time Fourier transform, the obtained frequency domain signal includes 1872 real sequences with a length of 257 and 1872 imaginary sequences with a length of 257.
[0130] Step 206: Perform model training on the first sub-model according to the first loss determined based on the frequency domain coding feature and the frequency domain transformation sequence corresponding to the clean audio signal to obtain a frequency domain processing sub-model.
[0131] In order to enable the first sub-model to learn the frequency domain characteristics of the clean audio signal in the sample audio signal, that is, the amplitude information and phase information of the signal, which are jointly determined by the real part and the imaginary part, the computer device can construct a loss function, that is, the first loss, based on the frequency domain coding characteristics and the frequency domain transformation characteristics corresponding to the clean audio signal in the sample audio signal, and use the first loss to train the model and update the model parameters of the first sub-model. When the training stop condition is met, the first sub-model learns the ability to extract the frequency domain characteristics of the clean signal from the audio signal. The trained model can be called a frequency domain processing sub-model.
[0132] In one embodiment, the step of determining the first loss includes: performing frequency domain transform processing on the clean audio signal to obtain a corresponding frequency domain transform sequence, the frequency domain transform sequence includes a real part sequence and an imaginary part sequence; determining the first loss based on the difference between the real part sequence corresponding to the clean audio signal and the real part features in the frequency domain coding features corresponding to the sample audio signal, and the difference between the imaginary part sequence corresponding to the clean audio signal and the imaginary part features in the frequency domain coding features corresponding to the sample audio signal.
[0133] Specifically, the computer device may perform an inverse Fourier transform on the clean audio signal used to generate the sample audio signal to obtain the frequency domain transform features corresponding to the clean audio signal, including a real sequence and an imaginary sequence. The computer device may perform the following formulas for the real sequence y r and the imaginary part sequence y i Calculate the first loss together:
[0134]
[0135] Among them, y r represents the real part sequence of the frequency domain transformation features obtained by Fourier transforming the clean audio signal, y i represents the imaginary part sequence in the frequency domain transformation feature obtained by Fourier transforming the clean audio signal, Represents the real part of the frequency domain coding feature corresponding to the sample audio signal, f i w (x) represents the imaginary part feature in the frequency domain coding feature corresponding to the sample audio signal, the subscript r represents the real part of the complex number, and the subscript i represents the imaginary part of the complex number.
[0136] like Figure 3 FIG. 1 is a schematic diagram of a model framework of a frequency domain processing sub-model in an embodiment. Figure 3The frequency domain processing sub-model includes a Fourier forward transform module, a real part processing network and an imaginary part processing network. The real part processing network and the imaginary part processing network are complex-LSTM structures based on real and imaginary part operations. After the sample audio signal is input into the frequency domain processing sub-model, it is Fourier forward transformed through the Fourier forward transform module to obtain the real part sequence and imaginary part sequence of the frequency domain signal. Both the real part sequence and the imaginary part sequence are encoded by the complex-LSTM feature to obtain the real part attention and the imaginary part attention. Based on the real part attention and the imaginary part attention, the frequency domain coding features of the clean audio signal are mined from the real part sequence and the imaginary part sequence. The frequency domain coding features can obtain the time domain signal through the inverse Fourier transform.
[0137] In one embodiment, since the frequency domain processing sub-model obtained by training according to the first loss has the ability to mine the frequency domain features of the clean signal from the audio signal, the computer device can directly connect the frequency domain processing sub-model with the inverse Fourier transform module to obtain the audio noise reduction model. That is, when the original audio signal needs to be subjected to noise reduction processing, it is only necessary to input the original audio signal into the trained frequency domain processing sub-model, obtain the corresponding frequency domain coding features, and then perform inverse Fourier transform on the frequency domain coding features to obtain the noise reduction signal corresponding to the original audio signal.
[0138] Step 208, the frequency domain processing sub-model is connected to the time domain processing sub-model to be trained and then trained together to obtain an audio noise reduction model for performing noise reduction processing on the audio signal.
[0139] Among them, the time domain processing sub-model is a model for extracting clean signals from audio signals from a time domain perspective. The time domain processing sub-model can adopt a neural network model. In this embodiment, in order to obtain a noise reduction signal with better sound quality, the computer device will further learn the characteristics of the sample audio signal in the time domain. Specifically, after the computer device obtains the frequency domain processing sub-model through model training, the frequency domain processing sub-model and the time domain processing sub-model to be trained are further integrated for training to obtain a trained time domain processing sub-model and a frequency domain processing sub-model with updated model parameters. After the integrated training is completed, the updated frequency domain processing sub-model is connected to the trained time domain processing sub-model to obtain an audio noise reduction model.
[0140] That is to say, the first half of the audio denoising model is learned in the frequency domain, and in order to make full use of the amplitude and phase information of the audio signal, a complex-LSTM structure based on real and imaginary part operations is designed. The second half of the audio denoising model is further supplemented by learning in the time domain to obtain a denoised signal with better sound quality.
[0141] In one embodiment, when the computer device continues to integrate the frequency domain processing sub-model with the time domain processing sub-model to be trained, the loss of the frequency domain processing sub-model is no longer introduced. Instead, the model parameters of the time domain processing sub-model are updated only according to the loss of the time domain denoising sub-module, and the model parameters of the frequency domain processing sub-model are adjusted at the same time.
[0142] During the integrated training, the computer device can input the sample audio signal into the frequency domain processing sub-model obtained in the previous step, and after the frequency domain processing sub-model outputs the corresponding frequency domain coding features, the frequency domain coding features are inversely Fourier transformed to obtain a time domain signal, and then the time domain signal is input into the time domain processing sub-model, and the noise reduction signal is output through the processing of the time domain processing sub-model. The computer device can compare the clean audio signal in the sample audio signal with the noise reduction signal output by the time domain processing sub-model, and then calculate the loss function, that is, the loss function of this part of the time domain processing sub-model, and then perform gradient back propagation according to the loss function to adjust the model parameters of the time domain processing sub-model and the frequency domain processing sub-model. The loss function of this part of the time domain processing sub-model can be calculated using an indicator for evaluating the quality of the audio noise reduction effect, such as SNR (Signal Noise Ratio) or SI-SDR (Scale Invariant Signal-to-Distortion Ratio).
[0143] like Figure 4 FIG. 1 is a schematic diagram of an embodiment of cascading a frequency domain processing sub-model and a time domain processing sub-model to obtain an audio noise reduction model. Figure 4 During model training, the difference between the frequency domain coding features corresponding to the sample audio signal and the frequency domain transformation features corresponding to the clean audio signal that generates the sample audio signal is first used to construct the first loss to obtain the frequency domain processing sub-model. Then, the frequency domain processing sub-model is cascaded with the time domain processing sub-model to be trained, and the model is trained together according to the difference between the clean audio signal and the noise reduction signal output by the time domain processing sub-model. After the joint training is completed, in the obtained audio noise reduction model, the frequency domain processing sub-model and the time domain processing sub-model can make full use of the frequency domain information and time domain information of the audio signal, thereby obtaining a very good audio noise reduction effect.
[0144] The processing method of the above-mentioned audio noise reduction model, the audio noise reduction model includes a frequency domain processing sub-model and a time domain processing sub-model, wherein the frequency domain processing sub-model includes a real part processing network and an imaginary part processing network. After obtaining the original audio signal in the time domain, the real part processing network and the imaginary part processing network in the frequency domain processing sub-model are used to perform feature encoding on the real part sequence and the imaginary part sequence of the sample audio signal respectively. Such a network structure can fully learn the frequency domain information of the original audio signal, that is, the amplitude information and phase information represented by the real part sequence and the imaginary part sequence, so that the real part attention and the imaginary part attention obtained by encoding can be used for the sample audio signal. More and more accurate attention is given to the clean audio signal in the sample audio signal. In this way, based on the frequency domain coding features corresponding to the sample audio signal obtained by the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, the frequency domain processing sub-model is trained according to the first loss determined by the frequency domain coding features and the frequency domain transform sequence corresponding to the clean audio signal that generates the sample audio signal. The frequency domain processing sub-model can accurately learn the frequency domain features of the clean signal in the sample audio signal. Subsequently, the frequency domain processing sub-model is jointly trained with the time domain processing sub-model, and the denoising effect of the audio denoising model obtained will be better.
[0145] like Figure 5 As shown, in one embodiment, the real part processing network and the imaginary part processing network in the first sub-model based on the neural network are used to respectively perform feature encoding on the real part sequence and the imaginary part sequence obtained after the sample audio signal is transformed into the frequency domain signal, so as to obtain the real part attention and the imaginary part attention corresponding to the sample audio signal, including:
[0146] Step 502: Input the sample audio signal into the first sub-model based on the neural network.
[0147] Specifically, the computer device may set the model structure of the first sub-model in advance, call the first sub-model, and input the acquired sample audio signal into the first sub-model.
[0148] Step 504: In the first sub-model, frequency domain transform is performed on the sample audio signal to obtain a real part sequence and an imaginary part sequence corresponding to the sample audio signal.
[0149] Specifically, the computer device may perform frequency domain transformation on the sample audio signal in the first sub-model and then convert it into a frequency domain signal. For example, the sample audio signal may be subjected to short-time Fourier transformation to obtain a real part sequence and an imaginary part sequence in the frequency domain signal.
[0150] Step 506, feature encoding is performed on the real part sequence and the imaginary part sequence respectively through the real part processing network in the first sub-model to obtain the real part first encoding feature and the imaginary part first encoding feature.
[0151] Step 508, feature encoding is performed on the real part sequence and the imaginary part sequence respectively through the imaginary part processing network in the first sub-model to obtain the real part second encoding feature and the imaginary part second encoding feature.
[0152] In order to obtain a good noise reduction effect, the computer device learns the frequency domain features of the audio signal from a frequency domain perspective. After converting the time domain signal into a frequency domain signal, in order to fully learn the amplitude information and phase information of the audio signal, the computer device performs feature encoding on the real part sequence and the imaginary part sequence, and mines the frequency domain features of the clean audio signal from the time domain signal. Specifically, within the real part processing network and the imaginary part processing network, the computer device operates the input real part sequence and imaginary part sequence with the network internal parameter matrix to obtain the corresponding coding features.
[0153] Refer to the complex multiplication formula:
[0154] Complex number complex1 = a + jb;
[0155] Complex number complex2 = c + jd;
[0156] Multiply complex number complex1 by complex number complex2:
[0157] complex1·complex2=(a·cb·d)+j(a·d+b·c);
[0158] The embodiment of the present invention defines a formula for the real part processing network and the imaginary part processing network to operate on the real part sequence and the imaginary part sequence:
[0159] L rr =LSTM r (X r );L ir =LSTM r (X i );
[0160] L ri =LSTM i (X r );L ii =LSTM i (X i );
[0161] L out =L rr -L ii +j(L ri +L ir );
[0162] Among them, X r It means that the real part sequence is obtained after the input sample audio signal X is transformed into the frequency domain, X irepresents the imaginary part sequence obtained after frequency domain transformation of the input sample audio signal X; L rr Represents the processed real part processing network LSTM r X r The result of the operation after processing is the first encoding feature L of the real part rr ; L ir Represents the real part processing network LSTM r X i The result of the operation after processing is the imaginary first coding feature L ir ; L ri Represents the imaginary part processing network LSTM i X r The result of the operation after processing is the real second coding feature; L ii Indicates the imaginary part processing network LSTM i X i The result of the operation after the processing is the imaginary second coding feature; L out Indicates the calculation results obtained after each layer of real part processing network and imaginary part processing network, including L rr -L ii The real part of L ri +L ir The imaginary part of .
[0163] That is to say, the output results of each layer in complex-LSTM are also divided into real and imaginary parts, where the real part is related to both the real part sequence and the imaginary part sequence in the frequency domain signal corresponding to the sample audio signal, and the imaginary part is related to both the real part sequence and the imaginary part sequence in the frequency domain signal corresponding to the sample audio signal.
[0164] like Figure 6 FIG. 1 is a schematic diagram showing a real part processing network and an imaginary part processing network performing feature encoding on a real part sequence and an imaginary part sequence in an embodiment. Figure 6 , the real part sequence X r With the imaginary part sequence X i , the network parameters W of the network that will be processed by the real part r Network parameters for processing networks with imaginary parts - W i Together, they are mapped into real attention (i.e., L rr -L ii ), the real part sequence X r With the imaginary part sequence X i , the network parameters of the network that will be processed by the real part - W r The network parameters W of the imaginary part processing network i Together, they are mapped into the imaginary attention (i.e., L ri +L ir ).
[0165] Step 510, obtain the real part attention corresponding to the sample audio signal according to the real part first coding feature and the imaginary part second coding feature, and obtain the imaginary part attention corresponding to the sample audio signal according to the real part second coding feature and the imaginary part first coding feature.
[0166] According to the formula defined above: L out =L rr -L ii +j(L ri +L ir );
[0167] The computer device can process the real part of the first encoding feature L output by the real part processing network rr The imaginary second encoding feature L output by the imaginary processing network ii The difference is used as the real part attention corresponding to the sample audio signal, and the real second encoding feature L output by the imaginary processing network is ri The imaginary part of the real part processing network output is the first encoded feature L ir The sum of is taken as the imaginary attention corresponding to the sample audio signal.
[0168] In this way, after the feature encoding of the real part processing network and the imaginary part processing network in the first sub-model, the obtained real part attention refers to the real part and imaginary part of the sample audio signal, and the obtained imaginary part attention specifically refers to the real part and imaginary part of the sample audio signal, which can make full use of the multi-faceted information of the sample audio information and provide explainability for the subsequent better noise reduction effect.
[0169] In one embodiment, the real part processing network and the imaginary part processing network in the first sub-model include at least two layers. The computer device obtains the complex number results output by the real part processing network and the imaginary part processing network of the previous layer, splits them into real and imaginary parts and uses them as the input of the current layer. The real part processing network and the imaginary part processing network of the current layer are used to perform feature encoding respectively to obtain each encoded feature, and each encoded feature is calculated according to the above formula to obtain the complex number result output by the current layer, which is input to the next layer for the same feature encoding and calculation, and so on, until the complex number result output by the last layer is obtained, and it is split into real and imaginary parts and used as real attention and imaginary attention respectively.
[0170] In some embodiments, the complex result output by the last layer may also be processed by the fully connected layer to obtain the final real attention and imaginary attention. The fully connected layer is used to perform matrix multiplication processing on its input features and the network parameters corresponding to the fully connected layer, thereby outputting the corresponding features. Specifically, the real part processing network of the last layer is connected to the first fully connected layer, and the imaginary part processing network of the last layer is connected to the second fully connected layer, that is, the real part of the complex result output by the real part processing network and the imaginary part processing network of the last layer is the input of the first fully connected layer, and the imaginary part of the complex result output by the real part processing network and the imaginary part processing network of the last layer is the input of the second fully connected layer. The first fully connected layer can be used to perform matrix multiplication processing on the real part and the network parameters corresponding to the first fully connected layer to obtain the real attention corresponding to the original audio signal, and the second fully connected layer can be used to perform matrix multiplication processing on the imaginary part and the network parameters corresponding to the second fully connected layer to obtain the imaginary attention corresponding to the original audio signal.
[0171] It should be noted that the frequency domain processing subnetwork includes at least two layers of real part processing networks and imaginary part processing networks, wherein the real part processing network and imaginary part processing network of the first layer are used to receive the real part sequence and imaginary part sequence corresponding to the original audio signal, the real part processing network and imaginary part processing network of the last layer are used to output the real part attention and imaginary part attention corresponding to the original audio signal, the real part processing network and imaginary part processing network of the "current layer" are used to describe the layer currently performing feature encoding in the frequency domain processing subnetwork, and the real part processing network and imaginary part processing network of the "previous layer" are used to describe the layer above the current layer in the frequency domain processing subnetwork, and the input data of the real part processing network and imaginary part processing network of the current layer are the output data of the real part processing network and imaginary part processing network of the previous layer. Moreover, the "current layer" is a relatively changing concept. For example, after the output result of the current layer s is obtained by encoding the features of the current layer s using the output result of the previous layer s-1, the output result of the current layer s is input to the next layer s+1. At this time, the next layer s+1 uses the output result of the current layer s for feature encoding. Then the next layer s+1 can be used as the new "current layer" and the current layer s can be used as the new previous layer. For example, the current layer can be the first layer or the last layer; the previous layer can be the first layer or the second layer; the next layer can be the second layer or the last layer.
[0172] It can be understood that in the embodiments of the present application, the current layer, the previous layer, and the last layer are not limited to the deployment positions in the frequency domain processing subnetwork, but are related to the processing order of data in the frequency domain processing subnetwork. For example, in the frequency domain processing subnetwork, when the data is processed from left to right, the previous layer can be set to the left of the next layer; when the data is processed from right to left, the previous layer can also be set to the right of the next layer; when the data is processed from top to bottom, the previous layer can also be set above the next layer; when the data is processed from bottom to top, the previous layer can also be set below the next layer. Similarly, the deployment positions of the last layer and the first layer in the frequency domain processing subnetwork are also related to the processing order in the frequency domain processing subnetwork. Refer to Figure 7 ,The data is processed from top to bottom in the frequency domain processing subnetwork. The upper layer is set on top of the lower layer, the first layer is set to the first layer, and the last layer is the last layer, that is, the second layer.
[0173] like Figure 7 FIG. 1 is a schematic diagram of a two-layer real part processing network and an imaginary part processing network in one embodiment. Figure 7 The real part sequence and imaginary part sequence in the frequency domain signal corresponding to the sample audio signal are input into the real part processing network and the imaginary part processing network of the first layer to obtain the complex result output by the first layer. The real part and sequence in the complex result are input into the real part processing network and the imaginary part processing network of the second layer to obtain the complex result output by the second layer. The real part of the complex result output by the second layer is input into the first fully connected layer to output the real part attention corresponding to the original audio signal. The imaginary part of the complex result output by the second layer is input into the second fully connected layer to output the imaginary part attention corresponding to the sample audio signal.
[0174] After the computer device obtains the real attention and the imaginary attention through the real processing network and the imaginary processing network based on real and imaginary part operations in the frequency domain processing sub-model, the frequency domain coding features corresponding to the sample audio signal can be represented by multiplying the attention by the original signal.
[0175] In one embodiment, based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, the frequency domain coding features corresponding to the sample audio signal are obtained, including: multiplying the real part sequence with the real part attention to obtain the real part of the frequency domain coding features corresponding to the sample audio signal; multiplying the imaginary part sequence with the imaginary part attention to obtain the imaginary part of the frequency domain coding features corresponding to the sample audio signal.
[0176] Specifically, the attention obtained for the clean audio signal in the sample audio signal includes two parts: real part attention and imaginary part attention. The computer device can ignore the phase and multiply them in the form of real number multiplication. That is, the real part attention is directly multiplied with the real part sequence in the frequency domain signal corresponding to the sample audio signal, and the product result is used as the real part of the frequency domain coding feature corresponding to the sample audio signal, and the imaginary part attention is multiplied with the imaginary part sequence in the frequency domain signal corresponding to the sample audio signal, and the product result is used as the imaginary part of the frequency domain coding feature corresponding to the sample audio signal.
[0177] That is, it is expressed by the following formula:
[0178]
[0179] in, represents the frequency domain coding feature corresponding to the sample audio signal, X represents the real part sequence obtained after the input sample audio signal X is transformed in the frequency domain, X i represents the imaginary part sequence obtained after frequency domain transformation of the input sample audio signal X, represents the real part of attention, Represents the imaginary part of attention.
[0180] In one embodiment, based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, the frequency domain coding feature corresponding to the sample audio signal is obtained, including: multiplying the real part sequence and the real part attention to obtain a first result, multiplying the imaginary part sequence and the imaginary part attention to obtain a second result, and using the difference between the first result and the second result as the real part of the frequency domain coding feature corresponding to the sample audio signal; multiplying the real part sequence and the imaginary part attention to obtain a third result, multiplying the imaginary part sequence and the real part attention to obtain a fourth result, and using the sum of the third result and the fourth result as the imaginary part of the frequency domain coding feature corresponding to the sample audio signal.
[0181] Specifically, the attention obtained for the clean audio signal in the sample audio signal includes two parts: real part attention and imaginary part attention. The computer device can multiply the real part and the imaginary part in the format of the real part and the imaginary part, that is, according to the formula of complex multiplication:
[0182]
[0183] in, represents the frequency domain coding feature corresponding to the sample audio signal, X r It means that the real part sequence is obtained after the input sample audio signal X is transformed into the frequency domain, X i represents the imaginary part sequence obtained after frequency domain transformation of the input sample audio signal X, represents the real part of attention, Represents the imaginary part of attention.
[0184] In one embodiment, based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, the frequency domain coding features corresponding to the sample audio signal are obtained, including: based on the real part sequence and the imaginary part sequence, the original amplitude information and the original phase information of the sample audio signal are obtained; based on the real part attention and the imaginary part attention, the predicted amplitude information and the predicted phase information of the sample audio signal are obtained; according to the product of the original amplitude information and the predicted amplitude information, the amplitude information of the frequency domain coding features corresponding to the sample audio signal is obtained; according to the sum of the original phase information and the predicted phase information, the phase information of the frequency domain coding features corresponding to the sample audio signal is obtained.
[0185] Specifically, the attention obtained for the clean audio signal in the sample audio signal includes two parts: real part attention and imaginary part attention, and the computer device can multiply the amplitude and phase information. That is, according to the formula of amplitude and phase multiplication:
[0186]
[0187] X phase =arctan(X i / X r )
[0188]
[0189] in, represents the frequency domain coding feature corresponding to the sample audio signal, X mag Represents the real part sequence X based on the sample audio signal r With the imaginary part sequence X i The original amplitude information obtained, X phase Represents the real part sequence X based on the sample audio signal r With the imaginary part sequence X i The original phase information obtained is Represents real-part attention Attention to the imaginary part The predicted amplitude information of the sample audio signal obtained, Real-part attention Attention to the imaginary part The predicted phase information of the sample audio signal is obtained, and j represents a complex number.
[0190] In the above embodiment, the real part attention and the imaginary part attention of the output are multiplied with the real part sequence and the imaginary part sequence in the original signal respectively to obtain the final frequency domain coding features, so that the frequency domain processing sub-model gradually learns the ability of the output attention to obtain a clean signal when multiplied with the original signal.
[0191] Furthermore, since each machine learning task has some noise to a greater or lesser extent, for example, when training the sub-model corresponding to task A, it is easy to ignore the data-related noise and generalization performance. Since different tasks have different noise patterns, learning multiple tasks at the same time can obtain a more generalized representation. Multi-task learning refers to putting multiple related machine learning tasks together for learning. During the learning process, shallow shared representations are used to share and complement the information learned in related fields, and multiple tasks can promote each other. Therefore, in order to improve the generalization ability of the model, when training the time domain processing sub-model, the computer device can also set a noise scene classification task, and perform multi-task learning based on the time domain denoising task of the time domain processing sub-model and the noise scene classification task. In this embodiment, the noise scene classification task has a high correlation with the time domain denoising task, which can improve the effects of the two sub-tasks and improve the generalization ability of the overall model.
[0192] Noise scene classification has certain application significance. Users in different scenes have different sensitivities to noise, and different noise reduction levels can be used. For example, in the scene where users communicate with relatives and friends on a daily basis, appropriate noise suppression can meet the needs, and the degree of noise reduction can be slightly weaker; when users are in a multi-person conference call scene, users have higher requirements for eliminating noisy background sounds. Different degrees of noise reduction are performed according to the different noise scenes in which the users are located, which is conducive to improving the user experience. Then, after identifying the noise scene category in the audio, the noise reduction signal can be automatically output according to the corresponding noise reduction level, or the noise reduction signal can be output according to the noise reduction level input by the user.
[0193] In one embodiment, after obtaining the frequency domain processing sub-model, the computer device can set a second sub-model and a third sub-model connected to the frequency domain processing sub-model, wherein the second sub-model is used to learn the time domain characteristics of the signal to achieve noise reduction of the time domain signal, and the third sub-model is used to learn the time domain characteristics of the signal to achieve classification of the noise scene category of the time domain signal. The second sub-model and the third sub-model are both models based on neural networks. When multi-task learning is performed using the time domain noise reduction task and the noise scene classification task, the label information of the input sample audio signal also includes the noise scene category. As mentioned above, the computer device can mix the clean audio signal and the noise signal according to different signal-to-noise ratios to obtain a sample audio signal, then the label information of the sample audio signal also includes the noise scene category of the mixed noise signal, such as white noise, wind sound, subway sound, keyboard sound, mouse sound, and so on.
[0194] like Figure 8 As shown, in one embodiment, step 208, the frequency domain processing sub-model is connected with the time domain processing sub-model to be trained and then trained together to obtain an audio noise reduction model for performing noise reduction processing on the audio signal, including:
[0195] Step 802, connecting the frequency domain processing sub-model with the second sub-model and the third sub-model based on the neural network.
[0196] Specifically, the time domain noise reduction task and the noise scene classification task are two parallel sub-tasks. When performing multi-task learning, the computer device can connect the frequency domain processing sub-model with the second sub-model and the third sub-model respectively.
[0197] Step 804: transform the frequency domain coding features corresponding to the sample audio signal into a time domain signal.
[0198] Specifically, both the time domain noise reduction task and the noise scene classification task process the signal from the time domain perspective. Therefore, the computer device can transform the frequency domain coding features obtained by the frequency domain processing sub-model into time domain signals, and then input them into the second sub-model and the third sub-model respectively. The frequency domain coding features can be inverse Fourier transformed to obtain the time domain signal.
[0199] Optionally, after the computer device obtains the frequency domain coding features in the frequency domain processing sub-model, the frequency domain coding features are transformed into time domain signals through the frequency domain processing sub-model. The computer device can input the frequency domain coding features into the second sub-model and the third sub-model respectively, and each transforms the frequency domain coding features into time domain signals through inverse Fourier transform.
[0200] Step 806, construct a multi-task objective function based on the second loss determined by the denoised signal obtained by the second sub-model performing denoising on the time domain signal and the clean audio signal, and the third loss determined by the noise scene category obtained by the third sub-model performing noise classification on the time domain signal and the noise label category of the sample audio signal.
[0201] The computer device designs loss functions, namely the second loss and the third loss, respectively, according to the respective objectives of the audio noise reduction task and the noise scene classification task. The multi-task objective function is an objective function that integrates the losses of multiple tasks. The multi-task objective function in this embodiment integrates the second loss of the time domain noise reduction task and the third loss corresponding to the noise scene classification task. After integrating the second loss and the third loss, the same multi-task objective function is optimized to achieve the optimization of the two tasks.
[0202] In some embodiments, when the computer device continues to integrate the frequency domain processing sub-model with the second sub-model and the third sub-model, the constructed multi-task objective function no longer introduces the loss of the frequency domain processing sub-model, but only updates the model parameters of the second sub-model and the third sub-model according to the loss of the time domain denoising task corresponding to the second sub-model and the loss of the noise scene classification task corresponding to the third sub-model, and adjusts the model parameters of the frequency domain processing sub-model at the same time.
[0203] In some embodiments, in order to prevent the entire audio denoising model from being dominated by a certain task during the training process, resulting in poor final performance, the variance uncertainty is used to determine the weight of each task in the multi-task. The derivation process of the multi-task objective function is as follows:
[0204] Assuming the input of the model is x and the weight of the model is W, the output of the model is f w (x).
[0205] Then, for the regression task, a probability model of a Gaussian likelihood function can be defined:
[0206] p(y|f w (x)) = N(f w (x),σ 2 );
[0207] Among them, N represents the Gaussian likelihood function, and the mean of the Gaussian likelihood function is the output f of the model w (x), the standard deviation σ of the Gaussian likelihood function is also used as the noise of the model. According to the output f of the model w (x) is obtained by statistics.
[0208] For multiple tasks, the probability model of the multi-task likelihood function can be defined as:
[0209] p(y|f w (x))=p(y1|f w (x))...p(y k |f w (x));
[0210] Then, for the two regression tasks, the outputs are y1 and y2 respectively. The probability model of the multi-task objective function of the two regression tasks is defined according to the model parameters w and σ:
[0211]
[0212] Then we get the objective functions of multiple regression tasks:
[0213]
[0214] Among them, L1 and L2 represent the loss functions of the two regression tasks respectively. The goal of the model is to minimize this multi-task objective function L(W,σ1,σ2). From this objective function, it can be seen that when the model noise increases, the corresponding weight in the loss function will decrease. If the model noise decreases, the weight in the loss function will increase.
[0215] In the embodiment of the present application, the multi-task includes a time domain noise reduction task and a noise scene classification task. The former is a regression task and the latter is a classification task. The classification task usually connects the network output to the Softmax function to obtain a classification probability, and determines the final classification result according to the classification probability. That is, the probability model of the classification task is: p(y|f w (x)) = Softmax(f w (x));
[0216] Then, referring to the multi-task objective functions of the two regression tasks above, the multi-task objective functions of the time domain noise reduction task and the noise scene classification task in the embodiment of the present application can be determined as follows:
[0217]
[0218] Among them, L1 represents the second loss corresponding to the time domain denoising task, L2 represents the third loss corresponding to the noise scene classification task; σ1 represents the standard deviation obtained based on the output results of the second sub-model, and σ2 represents the standard deviation obtained based on the output results of the third sub-model.
[0219] In one embodiment, the steps of constructing the above-mentioned multi-task objective function include: obtaining multiple noise reduction signals obtained by performing noise reduction processing on time domain signals corresponding to multiple sample audio signals by the second sub-model, and determining the weight of the second loss according to the standard deviation of the multiple noise reduction signals; obtaining multiple noise scene categories obtained by performing noise classification on time domain signals corresponding to multiple sample audio signals by the third sub-model, and determining the weight of the third loss according to the standard deviation of the multiple noise scene categories; and fusing the second loss and the third loss according to their respective weights to obtain the multi-task objective function.
[0220] Based on the multi-task objective function derived above, the computer device can construct a multi-task objective function according to the second loss corresponding to the time domain noise reduction task of the second sub-model and the third loss corresponding to the noise scene classification task of the third sub-model. The second loss is determined based on the difference between the noise reduction signal output by the second sub-model and the clean audio signal for generating the sample audio signal, and the weight of the second loss is determined based on the standard deviation of the noise reduction signal corresponding to the multiple sample audio signals output by the second sub-model. The third loss is determined based on the difference between the noise scene category output by the third sub-model and the label category of the noise signal used to generate the sample audio signal, and the weight of the third loss is determined based on the standard deviation of the noise scene category corresponding to the multiple sample audio signals output by the third sub-model.
[0221] Step 808, training the frequency domain processing sub-model, the second sub-model and the third sub-model together according to the multi-task objective function to obtain an updated frequency domain processing sub-model, a trained time domain processing sub-model and a trained noise classification sub-model.
[0222] Specifically, when the computer device performs multi-task training, the sample audio signal is input into the frequency domain processing sub-model, and after the frequency domain processing sub-model outputs the corresponding frequency domain coding features, the frequency domain coding features are inverse Fourier transformed to obtain a time domain signal, and then the time domain signal is input into the second sub-model and the third sub-model, and the second sub-model processes and outputs a denoised signal, and the third sub-model outputs a noise scene category. The computer device can compare the clean audio signal in the sample audio signal with the denoised signal output by the second sub-model, and then calculate the second loss, calculate the third loss according to the noise label category of the noise signal in the sample audio signal and the noise scene category output by the third sub-model, construct a multi-task objective function according to the second loss and the third loss, and perform gradient back propagation according to the multi-task objective function, so as to train the second sub-model and the third sub-model, obtain a trained time domain processing sub-model and a trained noise classification sub-model, and update the model parameters of the frequency domain processing sub-model at the same time.
[0223] Step 810: Connect the updated frequency domain processing sub-model with the trained time domain processing sub-model to obtain an audio noise reduction model for performing noise reduction processing on the audio signal.
[0224] Specifically, the computer device may connect the updated frequency domain processing sub-model with the time domain processing sub-model obtained through multi-task learning as an audio noise reduction model. In other embodiments, the computer device may also connect the updated frequency domain processing sub-model with the time domain processing sub-model obtained through multi-task learning, and also connect it with the noise classification sub-model obtained through multi-task learning, to obtain an audio noise reduction model for simultaneously performing noise reduction processing on audio signals and noise scene classification.
[0225] In this embodiment, the noise scene classification task and the time domain noise reduction task are highly correlated. Multi-task learning can simultaneously improve the effects of the two subtasks, thereby improving the generalization ability of the overall model.
[0226] In one embodiment, the step of determining the second loss includes: encoding the time domain signal through the encoder in the second sub-model to obtain a time domain coding vector, extracting features from the time domain coding vector through the time series feature extraction network in the second sub-model to obtain hidden features corresponding to the time domain signal, decoding based on the time domain coding vector and the hidden features through the decoder in the second sub-model to obtain a noise reduction signal corresponding to the sample audio signal in the second sample set; and constructing the second loss based on the noise reduction signal and the clean audio signal.
[0227] Among them, the second sub-model adopts the convolution-based Encoder-Decoder structure. The Encoder-Decoder structure converts the input sequence into another sequence output. In this framework, the encoder converts the sequence corresponding to the input time domain signal into a vector, and the decoder accepts the vector and generates the output sequence in chronological order. The encoder and decoder can use the same type of neural network model or different types of neural network models. For example, the encoder and decoder can both be CNN (Convolutional Neural Networks) models, or the encoder can use the RNN (Recurrent Neural Networks) model and the decoder can use the CNN model. The time series feature extraction network of the second sub-model can use the LSTM model, for example, a 2-layer LSTM network can be used to learn the time series dependency of the input time domain signal.
[0228] Specifically, the computer device inputs the time domain signal output by the frequency domain sub-model into the second sub-model, and converts the input time domain signal into a time domain coding vector through the encoder of the second sub-model, and then passes through the intermediate time series feature extraction network to extract the intrinsic connection of the time domain coding vector in time series and mine hidden features. In this embodiment, the time series feature extraction network is an intermediate layer relative to the encoder as the input layer and the decoder as the output layer, also known as a hidden layer, so the features extracted by the time series feature extraction network are called hidden features. Finally, the decoder decodes based on the time domain coding vector and the hidden features, and outputs a noise reduction signal. The computer device constructs a second loss based on the noise reduction signal output by the second sub-model and the clean audio signal used to generate the sample audio signal.
[0229] As shown in Fig. 9(a), it is a schematic diagram of the structure of the second sub-model used for training the time domain processing sub-model in one embodiment. Referring to Fig. 9(a), the second sub-model adopts an Encoder-Decoder structure based on convolution, the encoder and decoder can be 1-dimensional convolution, and the middle layer of the second sub-model adopts an LSTM network.
[0230] In one embodiment, a second loss is constructed based on the noise reduction signal and the clean audio signal, including: projecting the noise reduction signal corresponding to the sample audio signal in the vertical direction and horizontal direction of the clean audio signal respectively to obtain a vertical projection vector and a horizontal projection vector; and obtaining a second loss based on the vertical projection vector and the horizontal projection vector.
[0231] Specifically, the computer device can project the vector corresponding to the output noise reduction signal to the vertical and horizontal directions of the direction of the vector corresponding to the clean audio signal, and then use the vertical projection vector as the denominator and the horizontal projection vector as the numerator, and then take the logarithm to obtain the second loss. It can be seen that when the noise reduction signal output by the second sub-model is parallel to the clean audio signal, the second loss is smaller, and when the vector direction deviation between the vector corresponding to the noise reduction signal and the vector corresponding to the clean audio signal is greater, the larger the vertical projection vector is, the greater the second loss value is.
[0232] In one embodiment, the second loss can be determined by the following formula:
[0233]
[0234] Y E =Y * -Y T ;
[0235]
[0236] Among them, Y E Represents the vertical projection vector, Y T Represents the horizontal projection vector, Y true Represents a clean audio signal, Y * Represents the denoised signal output by the second sub-model.
[0237] Referring to FIG9(b), a schematic diagram of the relationship between the second loss and the projection vector in one embodiment is shown. Referring to the right side of FIG9(b), when the noise reduction signal Y output by the second sub-model is * With clean audio signal Y true When parallel, the horizontal projection vector Y T Maximum, vertical projection vector Y E The smaller the second loss is, the larger the deviation of the vector direction between the vector corresponding to the noise reduction signal and the vector corresponding to the clean audio signal is, the larger the horizontal projection vector Y T The smaller the vertical projection vector Y E The larger it is, the greater the second loss value is.
[0238] In this embodiment, by projecting the vector corresponding to the noise reduction signal output by the second sub-model in the direction of the clean audio model, the difference between the input sample audio signal and the output noise reduction signal can be reasonably and effectively represented.
[0239] In one embodiment, the step of determining the third loss includes: encoding the time domain signal through the encoder in the third sub-model to obtain a time domain coding vector, extracting features from the time domain coding vector through the temporal feature extraction network in the third sub-model to obtain hidden features corresponding to the time domain signal, and predicting the noise scene category of the sample audio signal based on the hidden features through the output layer in the third sub-model; constructing the third loss according to the noise scene category and the noise label category of the noise signal used to generate the sample audio signal.
[0240] Specifically, the computer device inputs the time domain signal output by the frequency domain sub-model into the third sub-model, converts the input time domain signal into a time domain coding vector through the encoder of the third sub-model, and then passes through the intermediate time series feature extraction network to extract the intrinsic connection of the time domain coding vector in time series and mine hidden features. Finally, the noise scene category of the sample audio signal is output based on the hidden features through the output layer. The computer device constructs a third loss based on the noise scene category output by the third sub-model and the noise label category of the noise signal used to generate the sample audio signal. The time series feature extraction network of the third sub-model can use LSTM, for example, a 2-layer LSTM network.
[0241] The output layer of the third sub-model includes a fully connected layer, an activation layer, and a normalization layer. The fully connected layer in the output layer receives the hidden features extracted by the feature extraction network, performs matrix multiplication on the hidden features and the model parameters corresponding to the fully connected layer, maps the hidden features to the sample space, and finally, after the nonlinear characteristics are introduced through the activation layer, the noise scene category corresponding to the sample audio signal is output through the normalization layer (softmax function).
[0242] In one embodiment, the third loss corresponding to the noise scene classification task may adopt a cross entropy loss function, which is expressed by the following formula:
[0243]
[0244] Among them, K represents the total number of noise classification categories. For example, if noises from 5 different scenes are used to generate a large number of sample audio signals, the value of K is 5; yi represents the noise label category of the noise signal used to generate the currently processed sample audio signal. For example, if the true label is the i-th category, then yi=1, otherwise yi=0; pi represents the noise scene category of the sample audio signal predicted by the third sub-model, that is, the probability that the noise belongs to category i, which is calculated by the softmax function of the output layer.
[0245] like Fig.10 FIG. 1 is a schematic diagram of the structure of the third sub-model used for training the noise classification sub-model in one embodiment. Fig.10The input layer of the third sub-model is an encoder, and the middle layer of the third sub-model adopts a two-layer LSTM structure. The noise scene is classified and processed according to the noise information learned by this part of the structure.
[0246] like Fig.11 FIG. 1 is a schematic diagram of a model structure for performing multi-task learning on a time domain noise reduction task and a noise scene classification task in one embodiment. Fig.11 The sample audio signal is input into the frequency domain processing sub-model 1102, and is transformed into a frequency domain signal through the Fourier forward transform model in the frequency domain processing sub-model 1102. The frequency domain signal includes a real sequence and an imaginary sequence. The real sequence and the imaginary sequence are feature encoded respectively through the real processing network and the imaginary processing network based on complex-LSTM in the frequency domain processing sub-model to obtain real attention and imaginary attention. Based on the real attention, imaginary attention, real sequence and imaginary sequence, the frequency domain coding feature is obtained. The frequency domain processing sub-model is transformed into a time domain signal through the Fourier inverse transform module in the frequency domain processing sub-model, and is input into the time domain processing sub-model 1104 and the noise classification sub-model 1106 respectively. The encoder in the time domain processing sub-model 1104 is used to encode the time domain signal to obtain a time domain coding vector. The time domain coding vector is feature extracted by the LSTM-based temporal feature extraction network in the time domain processing sub-model 1104 to obtain hidden features corresponding to the time domain signal. The time domain coding vector is then input into the decoder in the time domain processing sub-model 1104. The decoder outputs a noise reduction signal corresponding to the sample audio signal. The output noise reduction signal and the clean audio signal in the sample audio signal can construct a second loss. The time domain signal is noise classified by the noise classification sub-model 1106. The output noise scene category and the noise label category of the noise signal in the sample audio signal can construct a third loss. Based on the second loss and the third loss, a multi-task objective function can be constructed for multi-task learning.
[0247] In one embodiment, before performing multi-task learning, the computer device can set a second sub-model connected to the frequency domain processing sub-model after obtaining the frequency domain processing sub-model, first pre-train the second sub-model to obtain the pre-trained time domain processing sub-model, and then introduce a third sub-model for classifying noise scene categories, and perform integrated training on the entire model based on the multi-task objective function of the pre-trained time domain processing sub-model and the third sub-model.
[0248] like Fig.12 As shown, in one embodiment, step 208, the frequency domain processing sub-model is connected with the time domain processing sub-model to be trained and then trained together to obtain an audio noise reduction model for performing noise reduction processing on the audio signal, including:
[0249] Step 1202, connecting the frequency domain processing sub-model with the second sub-model and the third sub-model based on the neural network.
[0250] Specifically, the time domain noise reduction task and the noise scene classification task are two parallel sub-tasks. When multi-task learning is required, the computer device can connect the frequency domain processing sub-model with the second sub-model and the third sub-model respectively.
[0251] Step 1204: transform the frequency domain coding features corresponding to the sample audio signal into a time domain signal.
[0252] Specifically, both the time domain noise reduction task and the noise scene classification task process the signal from the time domain perspective. Therefore, the computer device can transform the frequency domain coding features obtained by the frequency domain processing sub-model into a time domain signal. The computer device can perform an inverse Fourier transform on the frequency domain coding features to obtain a time domain signal.
[0253] Step 1206, based on the second loss determined by the denoised signal obtained by performing denoising processing on the time domain signal by the second sub-model and the clean audio signal, the frequency domain processing sub-model and the second sub-model are jointly trained to obtain an updated frequency domain processing sub-model and a trained time domain processing sub-model.
[0254] The computer device first introduces the frequency domain processing sub-model into the second sub-model to pre-train the second sub-model before multi-task learning. This allows the second sub-model to have better initial network parameters to speed up the efficiency of subsequent multi-task learning with the noise classification sub-task.
[0255] In some embodiments, when the computer device trains the frequency domain processing sub-model and the second sub-model together, it only performs gradient backpropagation based on the second loss of the time domain denoising task corresponding to the second sub-model to train the second sub-model to obtain a pre-trained time domain processing sub-model, and at the same time updates the model parameters of the frequency domain processing sub-model.
[0256] Step 1208: Input the sample audio signal into the updated frequency domain processing sub-model to obtain the corresponding frequency domain coding features, transform the frequency domain coding features into time domain signals and input them into the trained time domain processing sub-model and the third sub-model respectively.
[0257] After obtaining the time domain processing sub-model, for the sample audio signals subsequently used for multi-task learning, the computer device inputs the time domain signal output by the frequency domain processing sub-model into the pre-trained time domain processing sub-model and the third sub-model at the same time.
[0258] Step 1210, construct a multi-task objective function based on the second loss determined by the denoised signal obtained by denoising the time domain signal through the trained time domain processing sub-model and the clean audio signal, and the third loss determined by the noise scene category obtained by the third sub-model performing noise classification on the time domain signal and the noise label category of the sample audio signal.
[0259] The computer device constructs a multi-task objective function according to the second loss of the time domain denoising task corresponding to the pre-trained time domain processing sub-model and the third loss of the noise scene classification task corresponding to the third sub-model.
[0260] Step 1212, according to the multi-task objective function, the frequency domain processing sub-model, the trained time domain processing sub-model and the third sub-model are jointly trained to obtain an updated frequency domain processing sub-model, an updated time domain processing sub-model and a noise classification sub-model.
[0261] The computer device performs gradient back propagation according to the multi-task objective function to train the third sub-model to obtain a trained noise classification sub-model, and simultaneously updates the model parameters of the frequency domain processing sub-model and the time domain processing sub-model.
[0262] Step 1214, connecting the updated frequency domain processing sub-model with the updated time domain processing sub-model to obtain an audio noise reduction model for performing noise reduction processing on the audio signal.
[0263] Specifically, the computer device may connect the updated frequency domain processing sub-model with the time domain processing sub-model obtained through multi-task learning as an audio noise reduction model. In other embodiments, the computer device may also connect the updated frequency domain processing sub-model with the time domain processing sub-model obtained through multi-task learning, and also connect it with the noise classification sub-model obtained through multi-task learning, to obtain an audio noise reduction model for simultaneously performing noise reduction processing on audio signals and noise scene classification.
[0264] like Fig.13 FIG. 1 is a schematic diagram of the steps of training an audio noise reduction model in an embodiment. Fig.13The training of the model is divided into three steps. First, the frequency domain processing sub-model is learned. The input of the frequency domain processing sub-model is a noisy sample audio signal, and the output is the real attention and imaginary attention of the network output based on complex-LSTM; secondly, the pre-trained frequency domain processing sub-model is connected to the time domain processing sub-model. The input of the pre-trained frequency domain processing sub-model is a noisy sample audio signal, and the output of the time domain processing sub-model is a denoised signal. The SI-SNR is used as the loss function based on the clean audio signal and the denoised signal in the sample audio signal to train the time domain processing sub-model; finally, the noise classification sub-model is connected, and then the entire network is multi-task learned to obtain the entire audio denoising model, and then the performance of the entire audio denoising model on the test set is observed.
[0265] In specific tasks, the increase in the number of model parameters can improve the model's representation learning ability and the model's index effect to a certain extent, but due to various performance requirements, the model is often required to achieve the best index evaluation under limited parameters as much as possible. Since the network structure used in a single task is computationally intensive when completing the two subtasks of denoising and classification at the same time, it hinders the model deployment. In order to ensure the performance of the model, the computer equipment can use the distillation strategy for the audio denoising model to further improve the evaluation index of the audio denoising model in the audio time domain denoising task.
[0266] like Fig.14 FIG. 1 is a flow chart of a process for performing model distillation training on a time domain denoising sub-model in one embodiment. Fig.14 , including the following steps:
[0267] Step 1402, a teacher model for performing noise reduction processing on an audio signal is obtained according to the audio noise reduction model, and a student model for performing noise reduction processing on an audio signal is constructed according to the frequency domain processing sub-model and the lightweight time domain noise reduction network.
[0268] Specifically, the computer device can first set a teacher model with a large amount of network parameters, and obtain a trained teacher model under full data training according to the implementation method provided above. In addition, the computer device constructs a student model for denoising the audio signal based on the trained frequency domain processing sub-model and the lightweight time domain denoising network.
[0269] Compared with the time-domain processing sub-model in the teacher model, the lightweight time-domain denoising network is a network with a smaller size, fewer network parameters, and less computation. For example, the network parameters of the teacher model are about 10 times larger than those of the student model. The teacher model is an audio denoising model with a larger model size, more network parameters, and a larger computation. Under full data training, the teacher model has better evaluation indicators than the student model. Therefore, such a teacher model can be used to guide the student model's ability in audio time-domain denoising tasks.
[0270] Step 1404: input the sample audio signal into the teacher model, obtain the corresponding frequency domain coding features through the frequency domain processing sub-model in the teacher model, transform the frequency domain coding features into time domain signals, encode the time domain signals through the encoder in the time domain processing sub-model in the teacher model, obtain the first time domain coding vector corresponding to the sample audio signal, and obtain the first noise reduction signal corresponding to the sample audio signal based on the first time domain coding vector through the decoder in the time domain processing sub-model.
[0271] The first noise reduction signal is a prediction result obtained by the teacher model by performing noise reduction processing on the input sample audio signal, and the prediction result can be used as the annotation information of the student model to guide the learning of the student model. In addition, the first time domain coding vector is a coding vector obtained by encoding the time domain signal by the encoder in the time domain processing sub-model in the teacher model.
[0272] Step 1406: input the sample audio signal into the student model, obtain the corresponding frequency domain coding features through the frequency domain processing sub-model in the student model, transform the frequency domain coding features into time domain signals, encode the time domain signals through the encoder in the lightweight time domain denoising network, obtain the second time domain coding vector corresponding to the sample audio signal, and obtain the second denoised signal corresponding to the sample audio signal based on the second time domain coding vector through the decoder in the lightweight time domain denoising network.
[0273] The second denoised signal is a prediction result obtained by the student model performing denoising on the input sample audio signal. The second time domain coding vector is a coding vector obtained by encoding the time domain signal by the encoder in the lightweight time domain denoising network in the student model.
[0274] Step 1408, based on the model distillation loss determined based on the mean square error loss between the first time domain coding vector and the second time domain coding vector, the mean square error loss between the first denoised signal and the second denoised signal, and the data structure loss between the first denoised signal and the second denoised signal, the student model is trained according to the model distillation loss to obtain a lightweight audio denoising model.
[0275] In this embodiment, model distillation training refers to using the prediction results output by the trained and highly accurate teacher model to guide the training of the student model to achieve knowledge transfer. Therefore, the computer device can construct a model distillation loss function based on the second denoised signal output by the student model and the first denoised signal output by the trained teacher model, and use the model distillation loss function to update the parameters of the student model.
[0276] In order to improve the performance of the student model and further optimize the evaluation indicators, this embodiment aligns the teacher model and the student model at the encoding layer, decoding layer and output results with certain physical meanings.
[0277] First, the encoding layer and the decoding layer are aligned. The first time domain encoding vector and the second time domain encoding vector are encoding vectors output by the encoding layer, and the first noise reduction signal and the second noise reduction signal are the results output by the decoding layer. The mean square error loss (MSE) is used to make the feature distribution of the two models in the encoding layer and the decoding layer global approximation.
[0278] In addition, the time domain coding vectors and denoised signals output by the encoders and decoders of the student model and the teacher model are first normalized and then the mean square error loss is calculated. This can prevent the feature distribution of the two models from being affected by noise or outliers. This normalization operation has certain indicator improvements.
[0279] The normalization operation is shown in the following formula:
[0280]
[0281] Where Z represents the original vector, such as the time domain coding vector output by the encoding layer or the vector corresponding to the noise reduction signal output by the decoding layer, Z i represents the eigenvalue of the i-th dimension in the original vector, and c represents the dimension of the output original vector.
[0282] Then, the mean square error loss corresponding to the encoding layer and the decoding layer can be expressed by the following formula:
[0283]
[0284] Where, χ represents a set of sample audio signals in a batch of input sample audio signals, |χ| is the number of sample audio signals in the set, and x i Represents any sample audio signal in the set, t i In the time domain processing sub-model of the teacher model, the decoder is based on the sample audio signal x i Output noise reduction signal, s i Indicates that the decoder in the lightweight time-domain denoising network in the student model is based on the sample audio signal xi The output noise reduction signal, In the time domain processing sub-model of the teacher model, the encoder is based on the sample audio signal x i The first time-domain encoding vector output is The encoder of the lightweight time-domain denoising network in the student model is based on the sample audio signal x i The second time domain coding vector output; L mse Represents the mean square error loss between the output results of the encoding layer and decoding layer of the teacher model and the output results of the encoding layer and decoding layer of the student model.
[0285] In addition, this embodiment also makes the student model fit the audio output of the teacher model from the data structure level by comparing the internal connection of the output results of the teacher model and the student model at the data level. The data structure loss includes the distance loss and angle loss of the output results of the teacher model and the student model.
[0286] Among them, the distance loss can be expressed by the following formula:
[0287]
[0288] Among them, (x i ,x j ) is any pair of sample audio signals in a batch of sample audio signals input, 2 In a set consisting of two pairs of sample audio signals in a batch of input sample audio signals, |χ 2 | is the number of sample audio signals in the set, t i Represents the teacher model based on the sample audio signal x i Output noise reduction signal, t j Represents the teacher model based on the sample audio signal x j Output noise reduction signal, s i Represents the student model based on the sample audio signal x i Output noise reduction signal, s j Represents the student model based on the sample audio signal x j The output noise reduction signal, μ represents the normalization processing of the output data; l δ represents Huber loss; L D Represents the loss of the teacher model output and the student model output in terms of data distance.
[0289] Among them, the angle loss can be expressed by the following formula:
[0290] ψ A (t i ,t j ,t k )=cos∠ti t j t k = <e ij ,e kj >
[0291]
[0292] ψ A (s i ,s j ,s k )=cos∠s i s j s k = <e ij ,e kj >
[0293]
[0294] Among them, (x i ,x j , x k ) are any three sample audio signals in a batch of sample audio signals input, 3 In a set consisting of any three sample audio signals in a batch of sample audio signals input, | 3 | is the number of sample audio signals in the set, t i Represents the teacher model based on the sample audio signal x i Output noise reduction signal, t j Represents the teacher model based on the sample audio signal x j The output noise reduction signal, e ij Indicates t i ,t j The unit vector in the direction, e jk Indicates t j ,t k The unit vector in the direction, s i Represents the student model based on the sample audio signal x i Output noise reduction signal, s j Represents the student model based on the sample audio signal x j The output noise reduction signal, μ represents the normalization processing of the output data; l δ represents Huber loss; L A Represents the loss of the teacher model output and the student model output in terms of data.
[0295] The computer equipment can determine the model distillation loss L according to the following formula: KD :
[0296] L KD =λ mse·L mse +λ D ·L D +λ A ·L A ;
[0297] Among them, L mse is the mean square error loss of the output features of the encoding layer and the decoding layer, L D is the distance comparison loss between the output features of the teacher model and the student model, L A is the angle comparison loss between the output features of the teacher model and the student model, λ mse , D , A They represent the adjustable hyper parameters of the corresponding losses.
[0298] like Fig.15 FIG. 1 is a schematic diagram of a framework for performing model distillation training on a time domain denoising sub-model in one embodiment. Fig.15 , construct a teacher model according to the above embodiments and perform training to obtain a trained teacher model. When constructing the student model, directly use the frequency domain processing sub-model in the teacher model, and connect to the lightweight time domain denoising network, and combine the frequency domain processing sub-model with the lightweight time domain denoising network to obtain the student model. According to the mean square error loss between the features output by the encoding layer and decoding layer of the time domain denoising network and the data structure loss of the features output by the decoding layer of the teacher model and the student model, the lightweight time domain denoising network of the student model is subjected to model distillation training. After the student model has basically converged, the lightweight time domain denoising network of the student model can be fine-tuned using the multi-task learning training method mentioned above to finally obtain a trained student model.
[0299] In a specific embodiment, referring to Fig.11 The model structure and design parameters of the audio denoising model based on attention mechanism and multi-task learning are listed as follows:
[0300]
[0301]
[0302] The Fourier transform here uses short-time Fourier transform, which divides a longer time signal into shorter segments of the same length, and calculates the Fourier transform on each shorter segment, i.e., the Fourier spectrum. The sampling rate is the sampling rate of the audio signal. By sampling the audio signal at this sampling rate, a discrete audio sequence can be obtained. The length of the sample audio is the duration of the audio signal, for example, it can be 15s. The Fourier transform window length is the length of the shorter segments of the same length, for example, it is set to 512. The Fourier transform window overlap rate is used to determine the displacement step of the window. Batch Size is the number of sample audio signals in each batch input to the audio denoising model. The number of LSTM hidden units is the number of hidden units in each time step of each layer of the real part processing network and the imaginary part processing network in the frequency domain processing subnetwork, for example, 128 hidden units. The number of LSTM layers is the number of layers of the real part processing network and the imaginary part processing network in the frequency domain processing subnetwork, for example, 2 layers. The inactivation rate of the fully connected layer is the proportion of neurons in the fully connected layer that do not participate in the operation. For example, 25% of the neurons can be randomly selected not to participate in the operation. The number of convolutional layer channels is the depth of the convolutional layer in the time domain processing subnetwork and the noise classification subnetwork. The convolution kernel size refers to the size of the convolution kernel used by the encoder and decoder in the time domain processing subnetwork and the noise classification subnetwork.
[0303] The values of the parameters in the above table are for illustration only. Models obtained by those skilled in the art after making corresponding adjustments to the values of the parameters according to actual needs based on the various embodiments provided by this application also fall within the protection scope of this application.
[0304] The processing method of the above-mentioned audio denoising model provided in the embodiment of the present application adopts an end-to-end deep learning model, the input is the audio containing noise, and the output is the denoised audio and the current noise scene category. Based on the network structure design of deep learning, an audio frequency domain information learning network with an attention mechanism is proposed to better distinguish noise from audio; the first half of the model is to improve the ordinary LSTM structure into a complex-LSTM structure based on complex numbers, which can fully learn the frequency domain information of the audio signal, and the second half of the model is to further supplement the learning in the time domain; in addition, a method of audio denoising and audio classification combining attention mechanism and multi-task learning is proposed, which can obtain an end-to-end model that can simultaneously realize denoising and noise scene classification; in addition, it is proposed to determine the weight of the multi-task loss function by variance uncertainty; in addition, in order to further improve the model effect, the model distillation training method adopted improves the performance upper limit of the model without changing the order of magnitude of the parameters, and the model training process takes into account the model deployment, further improves the performance of the time domain denoising sub-model, and increases the engineering practicality of the model.
[0305] In one embodiment, Fig.16 As shown, an audio noise reduction method is provided, which is applied to Figure 1 The computer device (terminal 102 or server 104) in the example is used for explanation, and the following steps are included:
[0306] Step 1602: Acquire an original audio signal in the time domain.
[0307] The original audio signal is a sound signal to be subjected to noise reduction processing, for example, it may be an original voice signal during a voice call, an original sound signal during a song recording or a song performance, or an original recorded sound signal during a dialogue recording.
[0308] Since the original audio signal is in a variety of noisy environments when it is generated, the original audio signal is a signal that carries a noise signal. For example, during daily video and audio calls, the callers may be in a variety of noisy environments, causing the voice call signal to carry a variety of noises, such as: the sound of vehicles passing by on a noisy street, the messy background sound of multiple people talking in the cafeteria, the loud keyboard tapping sound and the "clicking" sound of the mouse in an office scene, etc. During the call, the user naturally hopes to maintain an undisturbed and high-quality call level. Then, through the audio noise reduction method provided in the embodiment of the present application, the computer device can perform noise reduction processing on the original audio signal to obtain a corresponding noise reduction signal.
[0309] Both time domain and frequency domain are basic properties of signals. Analyzing signals from different dimensions and solving problems from different angles can be called domains. The time domain reflects the correspondence between mathematical functions or physical signals and time. It is the feedback of the real world and the only objectively existing domain. The frequency domain is a coordinate system used to describe the characteristics of a signal in the frequency domain. It shows the amount of signal within a frequency range. It is a way of auxiliary thinking constructed from a mathematical perspective. Frequency domain signals can be expressed in complex form, including real and imaginary parts. Frequency domain signals can also be expressed in the form of amplitude and phase.
[0310] In an embodiment of the present application, the original audio signal acquired by the computer device is a time domain signal. The original audio signal may be a continuous time domain signal, indicating the change of the intensity of the audio with continuous time. The original audio signal may also be a discrete time domain signal, indicating the change of the intensity of the audio with the sampling point in time. For example, the original audio signal may be a 15s original audio signal, or a discrete sequence obtained after sampling it, for example, when the sampling rate is 16K, the original audio signal is a discrete sequence with a length of 240000. Optionally, after the computer device acquires the continuous original audio signal, it may subsequently input the trained frequency domain processing sub-model, and the frequency domain processing sub-model samples it according to the sampling rate and divides it into multiple discrete sub-sequences according to the preset window length. Optionally, the computer device may also first sample the continuous original audio signal and divide it into multiple discrete sub-sequences according to the preset window length, and then subsequently input the trained frequency domain processing sub-model.
[0311] In some embodiments, the computer device can collect audio signals in real time through a local audio collection device and use the collected audio signals as original audio signals, such as the above Figure 1 The user's voice signal collected by the terminal 102 through the microphone is used as the original audio signal. In other embodiments, the original audio signal can also be a signal transmitted from other computer devices.
[0312] Step 1604, through the real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model, feature encoding is performed on the real part sequence and the imaginary part sequence obtained after the original audio signal is transformed into the frequency domain signal, so as to obtain the real part attention and the imaginary part attention corresponding to the original audio signal.
[0313] Among them, the frequency domain processing sub-model is a model that is pre-trained and has the ability to mine the frequency domain features of clean audio signals in the original audio signal. The frequency domain processing sub-model can adopt a deep learning model based on a neural network, such as LSTM (Long short-term memory). LSTM is a recurrent neural network with a special structure that can learn the long-term dependencies of long sequence inputs. In this embodiment, the frequency domain processing sub-model can learn the intrinsic connection between each discrete sub-sequence in the original audio signal.
[0314] The frequency domain processing sub-model can be structurally divided according to the function. In this embodiment, in order to make full use of the amplitude and phase information of the original audio signal, the frequency domain processing sub-model includes a real part processing network and an imaginary part processing network. The real part processing network is designed to utilize the real part information of the original audio signal, and the imaginary part processing network is designed to utilize the imaginary part information of the original audio signal. The real part processing network and the imaginary part processing network can both be LSTM-based network structures, and the real part processing network and the imaginary part processing network can both include at least one layer of LSTM network structure, that is, the LSTM overall structure capable of processing complex frequency domain signals includes an LSTM for processing real part sequences and an LSTM for processing imaginary part sequences, which can be referred to as complex-LSTM.
[0315] The real part processing network can be used to obtain the real part attention corresponding to the original audio signal, and the imaginary part processing network can be used to obtain the imaginary part attention corresponding to the original audio signal. The real part attention can be used to reflect the attention to the clean signal in the real frequency domain features of the original audio signal, and the imaginary part attention can be used to reflect the attention to the clean signal in the imaginary frequency domain features of the original audio signal. Since the goal defined in the frequency domain processing sub-model during training is to multiply the real part features in the output results of the real part processing network and the imaginary part processing network with the real part sequence of the original audio signal to obtain the real part of the clean audio signal, and to multiply the imaginary part features with the imaginary part sequence of the original audio signal to obtain the imaginary part of the clean audio signal, based on such a structure equivalent to attention, the output results of the real part processing network and the imaginary part processing network can express more and more accurate attention to the clean audio signal in the original audio signal. Therefore, the output results of the real part processing network and the imaginary part processing network can be called real part attention and imaginary part attention.
[0316] Usually, after the input signal is input into the neural network, the network parameters in the network layer of the neural network will operate on the input signal to obtain the operation result. Each layer of the network will receive the operation result output by the previous layer of the network, and then after the operation of the current layer of the network, output the operation result of the current layer as the input of the next layer. In this embodiment, the real part processing network and the imaginary part processing network in the frequency domain processing sub-model include at least one layer.
[0317] Specifically, the computer device can input the acquired original audio signal into the trained frequency domain processing sub-model. In the frequency domain processing sub-model, the original audio signal is first converted into a frequency domain signal, and then the real sequence and imaginary sequence of the frequency domain signal are respectively input into the real processing network and the imaginary processing network. The real processing network and the imaginary processing network will each perform feature encoding on the real sequence and the imaginary sequence, that is, the network parameters in the real processing network and the imaginary processing network operate on the input real sequence and imaginary sequence, and the real attention and imaginary attention corresponding to the original audio signal are obtained according to the output results of the last layer of the real processing network and the imaginary processing network.
[0318] In one embodiment, the computer device may use Fourier transform in the frequency domain processing sub-model to transform the original audio signal into a frequency domain signal, where the frequency domain signal includes a real part sequence and an imaginary part sequence.
[0319] In one embodiment, the frequency domain processing sub-model can be trained together with the time domain processing sub-model described below, and the trained frequency domain processing sub-model is connected with the time domain processing sub-model as an audio noise reduction model. That is, the time domain signal output by the frequency domain processing sub-model will be input into the time domain processing sub-model for further noise reduction processing to obtain the final noise reduction signal.
[0320] In one embodiment, the computer device may perform deep learning on the model structure of the model in advance to obtain an initial model, and then perform model training on the initial model through a sample audio signal to obtain a frequency domain processing sub-model.
[0321] In one embodiment, the frequency domain processing sub-model may include a Fourier transform module according to its function. After the computer device inputs the original audio signal into the trained frequency domain processing sub-model, the original audio signal is subjected to a Fourier transform through the Fourier transform module to obtain a corresponding frequency domain signal, which includes a real sequence and an imaginary sequence. The Fourier transform may be a short-time Fourier transform (STFT).
[0322] Step 1606, based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, obtain the frequency domain coding features corresponding to the original audio signal.
[0323] Among them, the frequency domain coding feature is the characteristic of the audio signal reflected from the frequency domain. As mentioned above, the real attention and imaginary attention output by the real processing network and the real processing network in the frequency domain processing sub-model can give more and more accurate attention to the clean audio signal in the original audio signal. Then, the computer device can dig out more parts related to the clean audio signal from the real sequence and imaginary sequence of the frequency domain signal of the original audio signal based on the real attention and imaginary attention, and then obtain the frequency domain coding features corresponding to the original audio signal. It can be understood that the frequency domain coding features include real and imaginary parts.
[0324] Step 1608: transform the frequency domain coding features into a time domain signal, and perform noise reduction processing on the time domain signal through the trained time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal.
[0325] Among them, the time domain processing sub-model is a model that performs noise reduction processing on the signal from the perspective of the time domain. The time domain processing sub-model can adopt a neural network model. For example, it can be a model structure designed by combining a convolutional neural network and a long short-term memory neural network. The trained time domain processing sub-model has the ability to perform noise reduction processing on the time domain signal, that is, the input of the time domain processing sub-model is the time domain signal, and the output is the noise-reduced time domain signal.
[0326] Specifically, after the computer device obtains the frequency domain coding features corresponding to the original audio signal output by the frequency domain processing sub-model, it can perform inverse Fourier transform on the frequency domain coding features to obtain a time domain signal. In order to further improve the effect of the noise reduction processing, the computer device further performs noise reduction processing on the time domain signal through the trained time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal.
[0327] In some embodiments, the computer device can transform the frequency domain coding features into time domain signals through the frequency domain processing submodel. In this case, the frequency domain processing submodel includes an inverse Fourier transform module. The frequency domain processing submodel performs inverse Fourier transform on the frequency domain coding features through the inverse Fourier transform module, outputs the time domain signal as the input of the time domain processing submodel, and the time domain processing submodel further performs noise reduction processing from the time domain perspective. In some embodiments, the computer can transform the frequency domain coding features into time domain signals through the time domain processing submodel. In this case, the frequency domain processing submodel does not include an inverse Fourier transform module, the time domain processing submodel includes an inverse Fourier transform module, the frequency domain coding features output by the frequency domain processing submodel serve as the input of the time domain processing submodel, and the time domain processing submodel performs inverse Fourier transform on the frequency domain coding features through the inverse Fourier transform module to obtain the time domain signal.
[0328] In some embodiments, step 1608 may also be replaced by: transforming the frequency domain coding features into a time domain signal to obtain a noise reduction signal corresponding to the original audio signal. Since the frequency domain coding features represent the frequency domain features corresponding to the clean audio signal in the original audio signal to a certain extent, the time domain signal obtained by inverse Fourier transform represents the clean audio signal in the original audio signal to a certain extent, that is, the signal obtained after noise reduction processing. Based on this, the computer device can directly transform the frequency domain coding features into a time domain signal to obtain a noise reduction signal corresponding to the original audio signal, without the need to use a time domain processing sub-model to further perform noise reduction processing on the time domain signal, thereby improving the noise reduction processing effect.
[0329] Reference Fig.11 , when the original audio signal is subjected to denoising, the original audio signal is input into the frequency domain processing sub-model. Then, the original audio signal is transformed into a frequency domain signal through the Fourier transform module in the frequency domain processing sub-model to obtain a real sequence and an imaginary sequence. The real sequence and the imaginary sequence are feature encoded through the real processing network and the imaginary processing network based on real and imaginary operations in the frequency domain processing sub-model to obtain real attention and imaginary attention. Then, the frequency domain coding features corresponding to the original audio signal are obtained based on the real sequence and the imaginary sequence, the real attention and the imaginary attention through the attention module in the frequency domain processing sub-model. Finally, the frequency domain coding features are transformed into time domain signals through the inverse Fourier transform module in the frequency domain processing sub-model, and then input into the trained time domain processing sub-model. The time domain processing sub-model performs denoising on the signal from the time domain perspective, and outputs the final denoised signal.
[0330] In the above-mentioned audio denoising method, the trained frequency domain processing sub-model includes a real part processing network and an imaginary part processing network. After obtaining the original audio signal in the time domain, the real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model are used to perform feature encoding on the real part sequence and the imaginary part sequence of the original audio signal respectively. Such a network structure can make full use of the frequency domain information of the original audio signal, that is, the amplitude information and phase information represented by the real part sequence and the imaginary part sequence, so that the real part attention and the imaginary part attention obtained by encoding can give more and more accurate attention to the clean audio signal in the original audio signal. In this way, the frequency domain coding features corresponding to the original audio signal obtained based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention can accurately represent the frequency domain features of the clean audio signal in the original audio signal, and the time domain signal obtained according to the frequency domain coding features can also accurately express the clean audio signal in the original audio signal, and the denoising effect is better; in addition, by further using the time domain processing sub-model to perform denoising on the time domain signal, the sound quality of the time domain signal can be further improved, and the effect of the denoised signal obtained will also be better.
[0331] like Fig.17 As shown, in one embodiment, the real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model are used to perform feature encoding on the real part sequence and the imaginary part sequence obtained after the original audio signal is transformed into the frequency domain signal, respectively, to obtain the real part attention and the imaginary part attention corresponding to the original audio signal, including:
[0332] Step 1702: input the original audio signal into the trained frequency domain processing sub-model.
[0333] Specifically, the computer device may call the frequency domain processing sub-model and input the acquired original audio signal into the frequency domain processing sub-model.
[0334] Step 1704: In the frequency domain processing sub-model, frequency domain transformation is performed on the original audio signal to obtain a real part sequence and an imaginary part sequence corresponding to the original audio signal.
[0335] Specifically, the computer device can perform frequency domain transformation on the original audio signal in the frequency domain processing sub-model and then convert it into a frequency domain signal. For example, the original audio signal can be short-time Fourier transformed to obtain a real part sequence and an imaginary part sequence in the frequency domain signal.
[0336] Step 1706, feature encoding is performed on the real part sequence and the imaginary part sequence respectively through the real part processing network in the frequency domain processing sub-model to obtain the real part first encoding feature and the imaginary part first encoding feature.
[0337] Step 1708, feature encoding is performed on the real part sequence and the imaginary part sequence respectively through the imaginary part processing network in the frequency domain processing sub-model to obtain the real part second encoding feature and the imaginary part second encoding feature.
[0338] In order to obtain a good noise reduction effect, the computer equipment converts the time domain signal into a frequency domain signal and makes full use of the amplitude information and phase information of the time domain signal for feature encoding, and mines the frequency domain features of the clean audio signal from the time domain signal, that is, using the real part sequence and the imaginary part sequence for feature encoding. Specifically, inside the real part processing network and the imaginary part processing network, the input real part sequence and the imaginary part sequence are multiplied with the network internal parameter matrix.
[0339] Refer to the complex multiplication formula:
[0340] Complex number complex1 = a + jb;
[0341] Complex number complex2 = c + jd;
[0342] Multiply complex number complex1 by complex number complex2:
[0343] complex1·complex2=(a·cb·d)+j(a·d+b·c);
[0344] The embodiment of the present application defines a formula for the real part processing network and the imaginary part processing network to operate on the real part sequence and the imaginary part sequence:
[0345] L rr =LSTM r (X r );L ir =LSTM r (X i );
[0346] L ri =LSTM i (X r );L ii =LSTM i (X i );
[0347] L out =L rr -L ii +j(L ri +L ir );
[0348] Among them, X r It means that the real part sequence is obtained after the frequency domain transformation of the input original audio signal X, X i represents the imaginary part sequence obtained after frequency domain transformation of the input original audio signal X; L rr Represents the processed real part processing network LSTM r X r The result of the operation after processing is the first encoding feature L of the real part rr ; L ir Represents the real part processing network LSTM r The result of the operation after processing Xi is the imaginary first coding feature L ir ; L ri Represents the imaginary part processing network LSTM i X r The result of the operation after processing is the real second coding feature; L ii Indicates the imaginary part processing network LSTM i X i The result of the operation after the processing is the imaginary second coding feature; L out Represents the calculation result obtained after each layer of real part processing network and imaginary part processing network.
[0349] That is to say, the output results of each layer in complex-LSTM are also divided into real and imaginary parts, where the real part is related to both the real part sequence and the imaginary part sequence in the frequency domain signal corresponding to the original audio signal, and the imaginary part is related to both the real part sequence and the imaginary part sequence in the frequency domain signal corresponding to the original audio signal.
[0350] Step 1710, obtain the real part attention corresponding to the original audio signal according to the real part first coding feature and the imaginary part second coding feature, and obtain the imaginary part attention corresponding to the original audio signal according to the real part second coding feature and the imaginary part first coding feature.
[0351] According to the formula defined above: L out =L rr -L ii +j(L ri +L ir );
[0352] The computer device can process the real part of the first encoding feature L output by the real part processing network rr The imaginary second encoding feature L output by the imaginary processing network ii The difference is used as the real part attention corresponding to the original audio signal, and the real second encoding feature L output by the imaginary part processing network is ri The imaginary part of the real part processing network output is the first encoded feature L ir The sum of is taken as the imaginary attention corresponding to the original audio signal.
[0353] In this way, after the feature encoding of the real part processing network and the imaginary part processing network in the frequency domain processing sub-model, the obtained real part attention refers to the real part and imaginary part of the original audio signal, and the obtained imaginary part attention specifically refers to the real part and imaginary part of the original audio signal, which can make full use of the multi-faceted information of the original audio information and provide explainability for the subsequent better noise reduction effect.
[0354] In one embodiment, the real part processing network and the imaginary part processing network in the frequency domain processing sub-model include at least two layers. The computer device obtains the complex number results output by the real part processing network and the imaginary part processing network of the previous layer, splits them into real and imaginary parts and uses them as the input of the current layer. The real part processing network and the imaginary part processing network of the current layer are used to perform feature encoding respectively to obtain each encoding feature, and each encoding feature is calculated according to the above formula to obtain the complex number result output by the current layer, which is input to the next layer for the same feature encoding and calculation, and so on, until the complex number result output by the last layer is obtained, and it is split into real and imaginary parts and used as real attention and imaginary attention respectively.
[0355] In some embodiments, the complex result output by the last layer may also be processed by the fully connected layer to obtain the final real attention and imaginary attention. The fully connected layer is used to perform matrix multiplication processing on its input features and the network parameters corresponding to the fully connected layer, thereby outputting the corresponding features. Specifically, the real part processing network of the last layer is connected to the first fully connected layer, and the imaginary part processing network of the last layer is connected to the second fully connected layer, that is, the real part of the complex result output by the real part processing network and the imaginary part processing network of the last layer is the input of the first fully connected layer, and the imaginary part of the complex result output by the real part processing network and the imaginary part processing network of the last layer is the input of the second fully connected layer. The first fully connected layer can be used to perform matrix multiplication processing on the real part and the network parameters corresponding to the first fully connected layer to obtain the real attention corresponding to the original audio signal, and the second fully connected layer can be used to perform matrix multiplication processing on the imaginary part and the network parameters corresponding to the second fully connected layer to obtain the imaginary attention corresponding to the original audio signal.
[0356] like Fig.18 As shown, in one embodiment, the real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model are used to perform feature encoding on the real part sequence and the imaginary part sequence obtained after the original audio signal is transformed into the frequency domain signal, respectively, to obtain the real part attention and the imaginary part attention corresponding to the original audio signal, including:
[0357] Step 1802, through the real part processing network of the first layer in the frequency domain processing sub-model, feature encoding is performed on the real part sequence and the imaginary part sequence respectively to obtain the real part first encoding feature and the imaginary part first encoding feature.
[0358] Step 1804, feature encoding is performed on the real part sequence and the imaginary part sequence respectively through the imaginary part processing network of the first layer in the frequency domain processing sub-model to obtain the real part second encoding feature and the imaginary part second encoding feature.
[0359] Step 1808, obtain the real part attention corresponding to the original audio signal at the first layer according to the real part first coding feature and the imaginary part second coding feature, and obtain the imaginary part attention corresponding to the original audio signal at the first layer according to the real part second coding feature and the imaginary part first coding feature.
[0360] Specifically, in the first layer of the real part processing network and the imaginary part processing network of the frequency domain processing subnetwork, the input is the real part sequence and the imaginary part sequence of the original audio signal, which are processed by the first layer of the real part processing network LSTM. r _1 and imaginary part processing network LSTM i _1 feature encoding, and obtain the complex results L including the output of the real part attention corresponding to the first layer and the imaginary part attention of the first layer out _1, that is:
[0361] L rr_1=LSTM r _1(X r );L ir _1=LSTM r _1(X i );
[0362] L ri _1=LSTM i _1(X r );L ii _1=LSTM i _1(X i );
[0363] L out _1=L rr _1-L ii _1+j(L ri _1+L ir _1).
[0364] Step 1810, iteratively perform feature encoding on the real attention and imaginary attention corresponding to the previous layer through the real part processing network and the imaginary part processing network of the current layer to obtain the real attention and imaginary attention corresponding to the current layer, and stop the iteration when the real attention and imaginary attention corresponding to the last layer are obtained.
[0365] Then in the second layer, the real part processing network LSTM r _2 and imaginary part processing network LSTM i _2, L out The real part of _1 is used as the real part sequence X of the second layer input r , L out The imaginary part of _1 is used as the imaginary part sequence X of the second layer input i :
[0366] L rr _2=LSTM r _2(X r );L ir _2=LSTM r _2(X i );
[0367] L ri _2=LSTM i _2(X r );L ii _1=LSTM i _2(X i );
[0368] L out _2=L rr _2-L ii _2+j(Lri _2+L ir _2).
[0369] L out _2 is the complex result of the second layer output, including the real part attention and the imaginary part attention output.
[0370] Similarly, when the real part processing network and the imaginary part processing network in the frequency domain processing sub-model include three layers, the real part processing network LSTM in the third layer is continued. r _3 and imaginary part processing network LSTM i _3, L out The real part of _2 is used as the real part sequence X of the third layer input r , L out The imaginary part of _2 is used as the imaginary part sequence X of the third layer input i , and similar processing is performed until the real attention and imaginary attention corresponding to the last layer are obtained, and the real attention and imaginary attention output by the last layer are used as the real attention and imaginary attention corresponding to the final original audio signal.
[0371] Each repetition of the above process in a layer of real part processing network and imaginary part processing network is called an "iteration" process. According to the above process, it is repeated multiple times, that is, multiple iterations are performed, and the real part attention and imaginary part attention output by the last layer are used as the real part attention and imaginary part attention corresponding to the final original audio signal.
[0372] After the computer device obtains the real attention and the imaginary attention through the real processing network and the imaginary processing network based on real and imaginary part operations in the frequency domain processing sub-model, the frequency domain coding features corresponding to the original audio signal can be expressed in the form of multiplying the attention by the original signal.
[0373] In one embodiment, based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, the frequency domain coding features corresponding to the original audio signal are obtained, including: multiplying the real part sequence with the real part attention to obtain the real part of the frequency domain coding features corresponding to the original audio signal; multiplying the imaginary part sequence with the imaginary part attention to obtain the imaginary part of the frequency domain coding features corresponding to the original audio signal.
[0374] Specifically, the attention obtained for the clean audio signal in the original audio signal includes two parts: real part attention and imaginary part attention. The computer device can ignore the phase and multiply them in the form of real number multiplication. That is, the real part attention is directly multiplied with the real part sequence in the frequency domain signal corresponding to the original audio signal, and the product result is used as the real part of the frequency domain coding feature corresponding to the original audio signal, and the imaginary part attention is multiplied with the imaginary part sequence in the frequency domain signal corresponding to the original audio signal, and the product result is used as the imaginary part of the frequency domain coding feature corresponding to the original audio signal.
[0375] That is, it is expressed by the following formula:
[0376]
[0377] in, Represents the frequency domain coding features corresponding to the original audio signal, X r It means that the real part sequence is obtained after the frequency domain transformation of the input original audio signal X, X i represents the imaginary part sequence obtained after frequency domain transformation of the input original audio signal X, represents the real part of attention, Represents the imaginary part of attention.
[0378] In one embodiment, based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, the frequency domain coding features corresponding to the original audio signal are obtained, including: multiplying the real part sequence and the real part attention to obtain a first result, multiplying the imaginary part sequence and the imaginary part attention to obtain a second result, and using the difference between the first result and the second result as the real part of the frequency domain coding features corresponding to the original audio signal; multiplying the real part sequence and the imaginary part attention to obtain a third result, multiplying the imaginary part sequence and the real part attention to obtain a fourth result, and using the sum of the third result and the fourth result as the imaginary part of the frequency domain coding features corresponding to the original audio signal.
[0379] Specifically, the attention obtained for the clean audio signal in the original audio signal includes two parts: real part attention and imaginary part attention. The computer device can multiply the real part and the imaginary part in the format of the real part and the imaginary part, that is, according to the formula of complex multiplication:
[0380]
[0381] in, Represents the frequency domain coding features corresponding to the original audio signal, X r It means that the real part sequence is obtained after the frequency domain transformation of the input original audio signal X, X i represents the imaginary part sequence obtained after frequency domain transformation of the input original audio signal X, represents the real part of attention, Represents the imaginary part of attention.
[0382] In one embodiment, frequency domain coding features corresponding to the original audio signal are obtained based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, including: obtaining the original amplitude information and the original phase information of the original audio signal based on the real part sequence and the imaginary part sequence, and obtaining the predicted amplitude information and the predicted phase information of the original audio signal based on the real part attention and the imaginary part attention; obtaining the amplitude information of the frequency domain coding features corresponding to the original audio signal according to the product of the original amplitude information and the predicted amplitude information; obtaining the phase information of the frequency domain coding features corresponding to the original audio signal according to the sum of the original phase information and the predicted phase information.
[0383] Specifically, the attention obtained for the clean audio signal in the original audio signal includes two parts: real part attention and imaginary part attention, and the computer device can multiply the amplitude and phase information. That is, according to the formula of amplitude and phase multiplication:
[0384]
[0385] X phase =arctan(X i / X r )
[0386]
[0387] in, Represents the frequency domain coding features corresponding to the original audio signal, X mag Represents the real part sequence X based on the original audio signal r With the imaginary part sequence X i The original amplitude information obtained, X phase Represents the real part sequence X based on the original audio signal r With the imaginary part sequence X i The original phase information obtained is Represents real-part attention Attention to the imaginary part The predicted amplitude information of the original audio signal is obtained. Real-part attention Attention to the imaginary part The predicted phase information of the original audio signal is obtained, and j represents a complex number.
[0388] In the above embodiment, the real part attention and the imaginary part attention of the output are multiplied with the real part sequence and the imaginary part sequence in the original signal respectively to obtain the final frequency domain coding features, so that the frequency domain processing sub-model gradually learns the ability of the output attention to obtain a clean signal when multiplied with the original signal.
[0389] In one embodiment, a time domain signal is subjected to noise reduction processing by a trained time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal, including: inputting the time domain signal into the trained time domain processing sub-model; encoding the time domain signal by an encoder in the time domain processing sub-model to obtain a time domain coding vector; extracting features from the time domain coding vector by a temporal feature extraction network in the time domain processing sub-model to obtain hidden features corresponding to the time domain signal; and decoding by a decoder in the time domain processing sub-model based on the time domain coding vector and the hidden features to obtain a noise reduction signal corresponding to the original audio signal.
[0390] Among them, the time domain processing sub-model adopts the Encoder-Decoder structure based on convolution. The Encoder-Decoder structure converts the input sequence into another sequence output. In this framework, the encoder converts the sequence corresponding to the input time domain signal into a vector, and the decoder accepts the vector and generates the output sequence in chronological order. The encoder and decoder can use the same type of neural network model or different types of neural network models. For example, the encoder and decoder can both be CNN (Convolutional Neural Networks) models, or the encoder can use the RNN (Recurrent Neural Networks) model and the decoder can use the CNN model. The time series feature extraction network of the time domain processing sub-model can use the LSTM model, for example, a 2-layer LSTM network can be used to learn the time series dependency of the input time domain signal.
[0391] Specifically, the computer device inputs the time domain signal output by the frequency domain sub-model into the time domain processing sub-model, and the encoder of the time domain processing sub-model converts the input time domain signal into a time domain coding vector, and then passes through the intermediate time series feature extraction network to extract the intrinsic connection of the time domain coding vector in time series and mine hidden features. In this embodiment, the time series feature extraction network is an intermediate layer relative to the encoder as the input layer and the decoder as the output layer, also called a hidden layer, so the features extracted by the time series feature extraction network are called hidden features. Finally, the decoder decodes based on the time domain coding vector and the hidden features to output a noise reduction signal.
[0392] In one embodiment, the method further includes: classifying the time domain signal through a trained noise classification sub-model to obtain the noise scene category of the original audio signal, wherein the noise classification sub-model is obtained by jointly training the model with the time domain processing sub-model.
[0393] Specifically, the computer device inputs the time domain signal output by the frequency domain sub-model into the noise classification sub-model, and outputs the noise scene category of the noise signal in the original audio signal through the noise classification sub-model.
[0394] In one embodiment, a time domain signal is classified and processed by a trained noise classification sub-model to obtain the noise scene category of the original audio signal, including: inputting the time domain signal into the trained noise classification sub-model; encoding the time domain signal through an encoder in the noise classification sub-model to obtain a time domain coding vector; extracting features from the time domain coding vector through a temporal feature extraction network in the noise classification sub-model to obtain hidden features corresponding to the time domain signal; and predicting the noise scene category of the original audio signal based on the hidden features through an output layer in the noise classification sub-model.
[0395] Specifically, the computer device inputs the time domain signal output by the frequency domain sub-model into the noise classification sub-model, and converts the input time domain signal into a time domain coding vector through the encoder of the noise classification sub-model, and then passes through the intermediate time series feature extraction network to extract the intrinsic connection of the time domain coding vector in time series and mine hidden features. Finally, the noise scene category of the original audio signal is output based on the hidden features through the output layer. The time series feature extraction network of the noise classification sub-model can use LSTM, for example, a 2-layer LSTM network.
[0396] The output layer of the noise classification sub-model includes a fully connected layer, an activation layer, and a normalization layer. The fully connected layer in the output layer receives the hidden features extracted by the feature extraction network, performs matrix multiplication on the hidden features and the model parameters corresponding to the fully connected layer, maps the hidden features to the sample space, and finally, after the nonlinear characteristics are introduced through the activation layer, the noise scene category corresponding to the original audio signal is output through the normalization layer (softmax function).
[0397] In one embodiment, the above method also includes: obtaining an input noise reduction level; determining a signal weight corresponding to the input noise reduction level, the signal weight including a first weight and a second weight respectively used to adjust the ratio between the original audio signal and the noise reduction signal; according to the first weight and the second weight, after fusing the original audio signal with the noise reduction signal, obtain an audio output signal corresponding to the input noise reduction level.
[0398] In this embodiment, the computer device can provide the user with a one-key noise reduction function, and provide the user with a plurality of noise reduction level options, so that the user can choose a suitable noise reduction level according to the different environments they are in. The computer device can not only output the noise reduction signal in real time, but also output whether it is in a noisy environment and what type of noise environment it is in.
[0399] In a specific embodiment, the audio noise reduction method comprises the following steps:
[0400] 1. Obtain the original audio signal in the time domain;
[0401] 2. Input the original audio signal into the trained frequency domain processing sub-model;
[0402] 3. In the frequency domain processing sub-model, the original audio signal is transformed into the frequency domain to obtain the real part sequence and the imaginary part sequence corresponding to the original audio signal;
[0403] 4. Through the real part processing network in the frequency domain processing sub-model, feature encoding is performed on the real part sequence and the imaginary part sequence respectively to obtain the real part first encoding feature and the imaginary part first encoding feature;
[0404] 5. Through the imaginary part processing network in the frequency domain processing sub-model, feature encoding is performed on the real part sequence and the imaginary part sequence respectively to obtain the real part second encoding feature and the imaginary part second encoding feature;
[0405] 6. Obtain the real part attention corresponding to the original audio signal according to the real part first coding feature and the imaginary part second coding feature;
[0406] 7. Obtain the imaginary part attention corresponding to the original audio signal according to the real part second coding feature and the imaginary part first coding feature;
[0407] 8. Multiply the real part sequence by the real part attention to obtain the real part of the frequency domain coding feature corresponding to the original audio signal;
[0408] 9. Multiply the imaginary part sequence by the imaginary part attention to obtain the imaginary part of the frequency domain coding feature corresponding to the original audio signal;
[0409] 10. Transform the frequency domain coding features into time domain signals;
[0410] 11. Input the time domain signal into the trained time domain processing sub-model;
[0411] 12. Encode the time domain signal through the encoder in the time domain processing sub-model to obtain a time domain coding vector;
[0412] 13. Through the time series feature extraction network in the time domain processing sub-model, the time domain coding vector is extracted to obtain the hidden features corresponding to the time domain signal;
[0413] 14. Decoding is performed based on the time domain coding vector and the hidden features through the decoder in the time domain processing sub-model to obtain the noise reduction signal corresponding to the original audio signal;
[0414] 15. Input the time domain signal into the trained noise classification sub-model;
[0415] 16. Encode the time domain signal through the encoder in the noise classification sub-model to obtain a time domain coding vector;
[0416] 17. Through the time series feature extraction network in the noise classification sub-model, the time domain coding vector is extracted to obtain the hidden features corresponding to the time domain signal;
[0417] 18. Predict the noise scene category of the original audio signal based on the hidden features through the output layer in the noise classification sub-model;
[0418] 19. Get the input noise reduction level;
[0419] 20. Determine a signal weight corresponding to the input noise reduction level, the signal weight comprising a first weight and a second weight respectively used to adjust a ratio between the original audio signal and the noise reduction signal;
[0420] 21. According to the first weight and the second weight, the original audio signal and the noise reduction signal are fused to obtain an audio output signal corresponding to the input noise reduction level.
[0421] In a specific application scenario, an instant messaging client is running on a computer device. When a user makes a voice call, the instant messaging client can collect the original voice signal. The original voice signal is input into a trained voice denoising model through the instant messaging client according to the method provided in the embodiment of the present application. In the frequency domain processing submodel of the voice denoising model, the real part sequence and the imaginary part sequence obtained after the original voice signal is transformed into the frequency domain signal are feature encoded through the real part processing network and the imaginary part processing network, respectively, and the real part attention and the imaginary part attention corresponding to the original voice signal are obtained. Based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, the frequency domain coding feature corresponding to the original voice signal is obtained. In the time domain processing submodel of the voice denoising model, the frequency domain coding feature is transformed into a time domain signal, and the time domain signal is subjected to denoising processing to obtain the denoised signal corresponding to the original voice signal, and then the denoised signal is transmitted to the other party of the call. Of course, the computer device can also first transmit the original audio signal to the call object, and then the computer device used by the call object uses the voice denoising method provided in the embodiment of the present application to denoise the original voice signal.
[0422] In addition, in the noise classification submodel of the speech noise reduction model, the frequency domain coding features can be transformed into time domain signals, and the time domain signals can be noise classified to obtain the noise scene category corresponding to the original speech signal. In addition, a noise reduction level switching control can be set on the voice call interaction interface, and different controls correspond to different noise reduction levels. The computer device can respond to the trigger operation of the noise reduction level switching control, adjust the ratio between the original audio signal and the noise reduction signal, and then output the mixed signal, which can perform different degrees of noise reduction and improve the user experience.
[0423] It should be understood that, although the various steps in the above flowchart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0424] In one embodiment, Fig.19 As shown, a processing device 1900 for an audio noise reduction model is provided. The device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The device specifically includes: an acquisition module 1902, a frequency domain coding training module 1904 and an integrated training module 1906, wherein:
[0425] An acquisition module 1902 is used to acquire a sample audio signal, where the sample audio information is generated based on a clean audio signal;
[0426] The frequency domain coding training module 1904 is used to perform feature coding on the real part sequence and the imaginary part sequence obtained after transforming the sample audio signal into the frequency domain signal through the real part processing network and the imaginary part processing network in the first sub-model based on the neural network, respectively, to obtain the real part attention and the imaginary part attention corresponding to the sample audio signal, and obtain the frequency domain coding feature corresponding to the sample audio signal based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention; perform model training on the first sub-model according to the first loss determined based on the frequency domain coding feature and the frequency domain transformation sequence corresponding to the clean audio signal, to obtain the frequency domain processing sub-model;
[0427] The integrated training module 1906 is used to connect the frequency domain processing sub-model and the time domain processing sub-model to be trained and train them together to obtain an audio noise reduction model for performing noise reduction processing on the audio signal.
[0428] In one embodiment, the frequency domain coding training module 1904 is specifically used to input the sample audio signal into the first sub-model based on the neural network; in the first sub-model, the sample audio signal is transformed in the frequency domain to obtain the real part sequence and the imaginary part sequence corresponding to the sample audio signal; through the real part processing network in the first sub-model, the real part sequence and the imaginary part sequence are feature encoded respectively to obtain the real part first coding feature and the imaginary part first coding feature; through the imaginary part processing network in the first sub-model, the real part sequence and the imaginary part sequence are feature encoded respectively to obtain the real part second coding feature and the imaginary part second coding feature; according to the real part first coding feature and the imaginary part second coding feature, the real part attention corresponding to the sample audio signal is obtained, and according to the real part second coding feature and the imaginary part first coding feature, the imaginary part attention corresponding to the sample audio signal is obtained.
[0429] In one embodiment, the frequency domain coding training module 1904 is specifically used to multiply the real part sequence with the real part attention to obtain the real part of the frequency domain coding feature corresponding to the sample audio signal; and multiply the imaginary part sequence with the imaginary part attention to obtain the imaginary part of the frequency domain coding feature corresponding to the sample audio signal.
[0430] In one embodiment, the frequency domain coding training module 1904 is specifically used to multiply the real part sequence by the real part attention to obtain a first result, multiply the imaginary part sequence by the imaginary part attention to obtain a second result, and use the difference between the first result and the second result as the real part of the frequency domain coding feature corresponding to the sample audio signal; multiply the real part sequence by the imaginary part attention to obtain a third result, multiply the imaginary part sequence by the real part attention to obtain a fourth result, and use the sum of the third result and the fourth result as the imaginary part of the frequency domain coding feature corresponding to the sample audio signal.
[0431] In one embodiment, the frequency domain coding training module 1904 is specifically used to obtain the original amplitude information and the original phase information of the sample audio signal based on the real part sequence and the imaginary part sequence; obtain the predicted amplitude information and the predicted phase information of the sample audio signal based on the real part attention and the imaginary part attention; obtain the amplitude information of the frequency domain coding feature corresponding to the sample audio signal according to the product of the original amplitude information and the predicted amplitude information; obtain the phase information of the frequency domain coding feature corresponding to the sample audio signal according to the sum of the original phase information and the predicted phase information.
[0432] In one embodiment, the frequency domain coding training module 1904 is specifically used to perform frequency domain transform processing on the clean audio signal to obtain a corresponding frequency domain transform sequence, which includes a real part sequence and an imaginary part sequence; the first loss is determined based on the difference between the real part sequence corresponding to the clean audio signal and the real part features in the frequency domain coding features corresponding to the sample audio signal, and the difference between the imaginary part sequence corresponding to the clean audio signal and the imaginary part features in the frequency domain coding features corresponding to the sample audio signal.
[0433] In one embodiment, the integrated training module 1906 is also used to connect the frequency domain processing sub-model with the second sub-model and the third sub-model based on the neural network; transform the frequency domain coding features corresponding to the sample audio signal into a time domain signal; construct a multi-task objective function based on the second sub-model performing noise reduction processing on the time domain signal to obtain a denoised signal and a second loss determined by the clean audio signal, the third loss determined by the noise scene category obtained by the third sub-model performing noise classification on the time domain signal and the noise label category of the sample audio signal; perform model training on the frequency domain processing sub-model, the second sub-model and the third sub-model together according to the multi-task objective function to obtain an updated frequency domain processing sub-model, a trained time domain processing sub-model and a trained noise classification sub-model; connect the updated frequency domain processing sub-model with the trained time domain processing sub-model to obtain an audio denoising model for denoising the audio signal.
[0434] In one embodiment, the integrated training module 1906 is also used to obtain multiple noise reduction signals obtained by the second sub-model performing noise reduction processing on the time domain signals corresponding to the multiple sample audio signals, and determine the weight of the second loss according to the standard deviation of the multiple noise reduction signals; obtain multiple noise scene categories obtained by the third sub-model performing noise classification on the time domain signals corresponding to the multiple sample audio signals, and determine the weight of the third loss according to the standard deviation of the multiple noise scene categories; and fuse the second loss and the third loss according to their respective weights to obtain a multi-task objective function.
[0435] In one embodiment, the integrated training module 1906 is also used to encode the time domain signal through the encoder in the second sub-model to obtain the time domain coding vector, perform feature extraction on the time domain coding vector through the time series feature extraction network in the second sub-model to obtain hidden features corresponding to the time domain signal, and perform decoding based on the time domain coding vector and the hidden features through the decoder in the second sub-model to obtain the noise reduction signal corresponding to the sample audio signal in the second sample set; and construct a second loss based on the noise reduction signal and the clean audio signal.
[0436] In one embodiment, the integrated training module 1906 is also used to project the noise reduction signal corresponding to the sample audio signal in the vertical direction and horizontal direction of the clean audio signal respectively to obtain a vertical projection vector and a horizontal projection vector; and obtain a second loss based on the vertical projection vector and the horizontal projection vector.
[0437] In one embodiment, the integrated training module 1906 is used to encode the time domain signal through the encoder in the third sub-model to obtain the time domain coding vector, perform feature extraction on the time domain coding vector through the time series feature extraction network in the third sub-model to obtain the hidden features corresponding to the time domain signal, and predict the noise scene category of the sample audio signal based on the hidden features through the output layer in the third sub-model; construct a third loss according to the noise scene category and the noise label category of the noise signal used to generate the sample audio signal.
[0438] In one embodiment, the integrated training module 1906 is further used to connect the frequency domain processing sub-model with the second sub-model and the third sub-model based on the neural network; transform the frequency domain coding features corresponding to the sample audio signal into a time domain signal; perform model training on the frequency domain processing sub-model and the second sub-model based on the noise reduction signal obtained by performing noise reduction processing on the time domain signal by the second sub-model and the second loss determined by the clean audio signal to obtain an updated frequency domain processing sub-model and a trained time domain processing sub-model; input the sample audio signal into the updated frequency domain processing sub-model to obtain the corresponding frequency domain coding features, transform the frequency domain coding features into time domain signals, and then input them into the trained time domain processing sub-model and the third sub-model respectively. model; construct a multi-task objective function based on the second loss determined by the denoised signal obtained by denoising the time domain signal through the trained time domain processing sub-model and the clean audio signal, and the third loss determined by the noise scene category obtained by the third sub-model through noise classification of the time domain signal and the noise label category of the sample audio signal; according to the multi-task objective function, the frequency domain processing sub-model, the trained time domain processing sub-model and the third sub-model are jointly trained to obtain an updated frequency domain processing sub-model, an updated time domain processing sub-model and a noise classification sub-model; the updated frequency domain processing sub-model is connected with the updated time domain processing sub-model to obtain an audio denoising model for denoising the audio signal.
[0439] In one embodiment, the processing device 1900 of the above-mentioned audio denoising model also includes: a distillation training module, which is used to obtain a teacher model for denoising the audio signal according to the audio denoising model, and construct a student model for denoising the audio signal according to the frequency domain processing sub-model and the lightweight time domain denoising network; input the sample audio signal into the teacher model, obtain the corresponding frequency domain coding features through the frequency domain processing sub-model in the teacher model, transform the frequency domain coding features into a time domain signal, encode the time domain signal through the encoder in the time domain processing sub-model in the teacher model, obtain a first time domain coding vector corresponding to the sample audio signal, and obtain a first denoised signal corresponding to the sample audio signal based on the first time domain coding vector through the decoder in the time domain processing sub-model; A sample audio signal is input into the student model, and the corresponding frequency domain coding features are obtained through the frequency domain processing sub-model in the student model. After the frequency domain coding features are transformed into time domain signals, the time domain signal is encoded through the encoder in the lightweight time domain denoising network to obtain a second time domain coding vector corresponding to the sample audio signal. A second denoised signal corresponding to the sample audio signal is obtained based on the second time domain coding vector through the decoder in the lightweight time domain denoising network; a model distillation loss is determined based on the mean square error loss between the first time domain coding vector and the second time domain coding vector, the mean square error loss between the first denoised signal and the second denoised signal, and the data structure loss between the first denoised signal and the second denoised signal, and the student model is trained according to the model distillation loss to obtain a lightweight audio denoising model.
[0440] The processing device 1900 of the above-mentioned audio noise reduction model, the audio noise reduction model includes a frequency domain processing sub-model and a time domain processing sub-model, wherein the frequency domain processing sub-model includes a real part processing network and an imaginary part processing network. After obtaining the original audio signal in the time domain, the real part processing network and the imaginary part processing network in the frequency domain processing sub-model respectively perform feature encoding on the real part sequence and the imaginary part sequence of the sample audio signal. Such a network structure can fully learn the frequency domain information of the original audio signal, that is, the amplitude information and the phase information represented by the real part sequence and the imaginary part sequence, so that the real part attention and the imaginary part attention obtained by encoding can be used for the sample More and more accurate attention is given to the clean audio signal in the audio signal. In this way, the frequency domain coding features corresponding to the sample audio signal are obtained based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention. The frequency domain processing sub-model is trained according to the first loss determined by the frequency domain coding features and the frequency domain transformation sequence corresponding to the clean audio signal that generates the sample audio signal. The frequency domain processing sub-model can accurately learn the frequency domain features of the clean signal in the sample audio signal. Subsequently, the frequency domain processing sub-model is combined with the time domain processing sub-model for model training, and the noise reduction effect of the obtained audio denoising model will be better.
[0441] For the specific definition of the processing device 1900 of the audio noise reduction model, please refer to the definition of the processing method of the audio noise reduction model in the above text, which will not be repeated here. The various modules in the above-mentioned processing device of the audio noise reduction model can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0442] In one embodiment, Fig. 20 As shown, an audio noise reduction device 2000 is provided. The device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The device specifically includes: an acquisition module 2002, a frequency domain encoding module 2004 and a time domain noise reduction module 2006, wherein:
[0443] An acquisition module 2002 is used to acquire an original audio signal in the time domain;
[0444] The frequency domain coding module 2004 is used to perform feature coding on the real part sequence and the imaginary part sequence obtained after the original audio signal is transformed into the frequency domain signal through the real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model, respectively, to obtain the real part attention and the imaginary part attention corresponding to the original audio signal; based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, obtain the frequency domain coding feature corresponding to the original audio signal;
[0445] The time domain noise reduction module 2006 is used to transform the frequency domain coding features into a time domain signal, and perform noise reduction processing on the time domain signal through the trained time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal.
[0446] In one embodiment, the frequency domain coding module 2004 is also used to input the original audio signal into a trained frequency domain processing sub-model; in the frequency domain processing sub-model, the original audio signal is transformed in the frequency domain to obtain a real sequence and an imaginary sequence corresponding to the original audio signal; through the real processing network in the frequency domain processing sub-model, the real sequence and the imaginary sequence are feature encoded respectively to obtain a real first coding feature and an imaginary first coding feature; through the imaginary processing network in the frequency domain processing sub-model, the real sequence and the imaginary sequence are feature encoded respectively to obtain a real second coding feature and an imaginary second coding feature; according to the real first coding feature and the imaginary second coding feature, the real attention corresponding to the original audio signal is obtained, and according to the real second coding feature and the imaginary first coding feature, the imaginary attention corresponding to the original audio signal is obtained.
[0447] In one embodiment, the real part processing network and the imaginary part processing network in the frequency domain processing submodel include at least two layers; the frequency domain coding module 2004 is also used to perform feature encoding on the real part sequence and the imaginary part sequence respectively through the real part processing network of the first layer in the frequency domain processing submodel to obtain the real part first coding feature and the imaginary part first coding feature; perform feature encoding on the real part sequence and the imaginary part sequence respectively through the imaginary part processing network of the first layer in the frequency domain processing submodel to obtain the real part second coding feature and the imaginary part second coding feature; obtain the real part attention corresponding to the original audio signal in the first layer according to the real part first coding feature and the imaginary part second coding feature, and obtain the imaginary part attention corresponding to the original audio signal in the first layer according to the real part second coding feature and the imaginary part first coding feature; iteratively perform feature encoding on the real part attention and the imaginary part attention corresponding to the previous layer respectively through the real part processing network and the imaginary part processing network of the current layer to obtain the real part attention and the imaginary part attention corresponding to the current layer, until the iteration is stopped when the real part attention and the imaginary part attention corresponding to the last layer are obtained.
[0448] In one embodiment, the frequency domain coding module 2004 is also used to multiply the real part sequence with the real part attention to obtain the real part of the frequency domain coding feature corresponding to the original audio signal; and multiply the imaginary part sequence with the imaginary part attention to obtain the imaginary part of the frequency domain coding feature corresponding to the original audio signal.
[0449] In one embodiment, the frequency domain coding module 2004 is also used to multiply the real part sequence by the real part attention to obtain a first result, multiply the imaginary part sequence by the imaginary part attention to obtain a second result, and use the difference between the first result and the second result as the real part of the frequency domain coding feature corresponding to the original audio signal; multiply the real part sequence by the imaginary part attention to obtain a third result, multiply the imaginary part sequence by the real part attention to obtain a fourth result, and use the sum of the third result and the fourth result as the imaginary part of the frequency domain coding feature corresponding to the original audio signal.
[0450] In one embodiment, the frequency domain coding module 2004 is also used to obtain the original amplitude information and the original phase information of the original audio signal based on the real part sequence and the imaginary part sequence, and to obtain the predicted amplitude information and the predicted phase information of the original audio signal based on the real part attention and the imaginary part attention; to obtain the amplitude information of the frequency domain coding feature corresponding to the original audio signal according to the product of the original amplitude information and the predicted amplitude information; and to obtain the phase information of the frequency domain coding feature corresponding to the original audio signal according to the sum of the original phase information and the predicted phase information.
[0451] In one embodiment, the time domain noise reduction module 2006 is also used to input the time domain signal into a trained time domain processing sub-model; encode the time domain signal through the encoder in the time domain processing sub-model to obtain a time domain coding vector; extract features from the time domain coding vector through the timing feature extraction network in the time domain processing sub-model to obtain hidden features corresponding to the time domain signal; and decode based on the time domain coding vector and the hidden features through the decoder in the time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal.
[0452] In one embodiment, the audio noise reduction device 2000 further includes:
[0453] The noise classification module is used to classify the time domain signal through the trained noise classification sub-model to obtain the noise scene category of the original audio signal. The noise classification sub-model is obtained by jointly training the model with the time domain processing sub-model.
[0454] In the previous embodiment, the noise classification module is also used to input the time domain signal into the trained noise classification sub-model; encode the time domain signal through the encoder in the noise classification sub-model to obtain the time domain coding vector; extract features from the time domain coding vector through the time series feature extraction network in the noise classification sub-model to obtain hidden features corresponding to the time domain signal; and predict the noise scene category of the original audio signal based on the hidden features through the output layer in the noise classification sub-model.
[0455] In one embodiment, the audio noise reduction device 2000 further includes:
[0456] The noise reduction gear adjustment module is used to obtain the input noise reduction level; determine the signal weight corresponding to the input noise reduction level, the signal weight includes a first weight and a second weight respectively used to adjust the ratio between the original audio signal and the noise reduction signal; according to the first weight and the second weight, the original audio signal and the noise reduction signal are merged to obtain an audio output signal corresponding to the input noise reduction level.
[0457] In the above-mentioned audio noise reduction device 2000, the trained frequency domain processing sub-model includes a real part processing network and an imaginary part processing network. After obtaining the original audio signal in the time domain, the real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model are used to perform feature encoding on the real part sequence and the imaginary part sequence of the original audio signal respectively. Such a network structure can make full use of the frequency domain information of the original audio signal, that is, the amplitude information and phase information represented by the real part sequence and the imaginary part sequence, so that the real part attention and the imaginary part attention obtained by encoding can give more and more accurate attention to the clean audio signal in the original audio signal. In this way, the frequency domain coding features corresponding to the original audio signal obtained based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention can accurately represent the frequency domain features of the clean audio signal in the original audio signal, and the time domain signal obtained according to the frequency domain coding features can also accurately express the clean audio signal in the original audio signal, and the noise reduction effect is better; in addition, by further using the time domain processing sub-model to perform noise reduction processing on the time domain signal, the sound quality of the time domain signal can be further improved, and the effect of the obtained noise reduction signal will also be better.
[0458] For the specific definition of the audio noise reduction device 2000, please refer to the definition of the audio noise reduction method above, which will not be repeated here. Each module in the above-mentioned audio noise reduction device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0459] In one embodiment, a computer device is provided, which may be Figure 1 The internal structure diagram of the terminal or server in the Fig.21 As shown. The computer device includes a processor, a memory and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an audio noise reduction method and / or a processing method for an audio noise reduction model is implemented.
[0460] Those skilled in the art will understand that Fig.21The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0461] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0462] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0463] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in the above-mentioned method embodiments.
[0464] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0465] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0466] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. An audio noise reduction method, characterized in that: The method comprises: Get the original audio signal in the time domain; By using the real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model, feature encoding is performed on the real part sequence and the imaginary part sequence obtained after the original audio signal is transformed into the frequency domain signal, and the real part attention and the imaginary part attention corresponding to the original audio signal are obtained, wherein the real part attention is used to reflect the attention to the clean signal in the real part frequency domain feature of the original audio signal, and the imaginary part attention is used to reflect the attention to the clean signal in the imaginary part frequency domain feature of the original audio signal; Based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, obtaining a frequency domain coding feature corresponding to the original audio signal; The frequency domain coding features are transformed into time domain signals, and the time domain signals are subjected to noise reduction processing through a trained time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal.
2. The method according to claim 1, characterized in that The real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model respectively perform feature encoding on the real part sequence and the imaginary part sequence obtained after the original audio signal is transformed into the frequency domain signal to obtain the real part attention and the imaginary part attention corresponding to the original audio signal, including: Inputting the original audio signal into the trained frequency domain processing sub-model; In the frequency domain processing sub-model, frequency domain transformation is performed on the original audio signal to obtain a real part sequence and an imaginary part sequence corresponding to the original audio signal; Through the real part processing network in the frequency domain processing sub-model, feature encoding is performed on the real part sequence and the imaginary part sequence respectively to obtain a real part first encoding feature and an imaginary part first encoding feature; Through the imaginary part processing network in the frequency domain processing sub-model, feature encoding is performed on the real part sequence and the imaginary part sequence respectively to obtain a real part second encoding feature and an imaginary part second encoding feature; According to the real part first coding feature and the imaginary part second coding feature, the real part attention corresponding to the original audio signal is obtained, and according to the real part second coding feature and the imaginary part first coding feature, the imaginary part attention corresponding to the original audio signal is obtained.
3. The method according to claim 1, characterized in that The real part processing network and the imaginary part processing network in the frequency domain processing sub-model include at least two layers; The real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model respectively perform feature encoding on the real part sequence and the imaginary part sequence obtained after the original audio signal is transformed into the frequency domain signal to obtain the real part attention and the imaginary part attention corresponding to the original audio signal, including: Through the real part processing network of the first layer in the frequency domain processing submodel, feature encoding is performed on the real part sequence and the imaginary part sequence respectively to obtain a real part first encoding feature and an imaginary part first encoding feature; Through the imaginary part processing network of the first layer in the frequency domain processing sub-model, feature encoding is performed on the real part sequence and the imaginary part sequence respectively to obtain a real part second encoding feature and an imaginary part second encoding feature; According to the real part first coding feature and the imaginary part second coding feature, obtain the real part attention corresponding to the original audio signal at the first layer, and according to the real part second coding feature and the imaginary part first coding feature, obtain the imaginary part attention corresponding to the original audio signal at the first layer; Iteratively encode the real attention and imaginary attention corresponding to the previous layer through the real part processing network and the imaginary part processing network of the current layer to obtain the real attention and imaginary attention corresponding to the current layer, and stop the iteration when the real attention and imaginary attention corresponding to the last layer are obtained.
4. The method according to claim 1, characterized in that: The obtaining, based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, a frequency domain coding feature corresponding to the original audio signal comprises: Multiplying the real part sequence by the real part attention to obtain the real part of the frequency domain coding feature corresponding to the original audio signal; The imaginary part sequence is multiplied by the imaginary part attention to obtain the imaginary part of the frequency domain coding feature corresponding to the original audio signal.
5. The method according to claim 1, characterized in that The obtaining, based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, a frequency domain coding feature corresponding to the original audio signal comprises: Multiplying the real part sequence by the real part attention to obtain a first result, multiplying the imaginary part sequence by the imaginary part attention to obtain a second result, and taking the difference between the first result and the second result as the real part of the frequency domain coding feature corresponding to the original audio signal; The real part sequence is multiplied by the imaginary part attention to obtain a third result, the imaginary part sequence is multiplied by the real part attention to obtain a fourth result, and the sum of the third result and the fourth result is used as the imaginary part of the frequency domain coding feature corresponding to the original audio signal.
6. The method according to claim 1, characterized in that The obtaining, based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, a frequency domain coding feature corresponding to the original audio signal comprises: Based on the real part sequence and the imaginary part sequence, original amplitude information and original phase information of the original audio signal are obtained, and based on the real part attention and the imaginary part attention, predicted amplitude information and predicted phase information of the original audio signal are obtained; Obtaining amplitude information of a frequency domain coding feature corresponding to the original audio signal according to the product of the original amplitude information and the predicted amplitude information; The phase information of the frequency domain coding feature corresponding to the original audio signal is obtained according to the sum of the original phase information and the predicted phase information.
7. The method according to claim 1, characterized in that The performing noise reduction processing on the time domain signal by using the trained time domain processing sub-model to obtain the noise reduction signal corresponding to the original audio signal includes: Inputting the time domain signal into a trained time domain processing sub-model; Encoding the time domain signal through the encoder in the time domain processing sub-model to obtain a time domain coding vector; Performing feature extraction on the time domain coding vector through the time series feature extraction network in the time domain processing sub-model to obtain hidden features corresponding to the time domain signal; The decoder in the time domain processing sub-model performs decoding based on the time domain coding vector and the hidden feature to obtain a noise reduction signal corresponding to the original audio signal.
8. The method according to claim 1, characterized in that The method further comprises: The time domain signal is classified and processed by using a trained noise classification sub-model to obtain the noise scene category of the original audio signal, wherein the noise classification sub-model is obtained by performing model training together with the time domain processing sub-model.
9. The method according to claim 8, characterized in that The classifying process of the time domain signal by using the trained noise classification sub-model to obtain the noise scene category of the original audio signal includes: Inputting the time domain signal into a trained noise classification sub-model; Encoding the time domain signal through the encoder in the noise classification sub-model to obtain a time domain coding vector; Performing feature extraction on the time domain coding vector through the time series feature extraction network in the noise classification sub-model to obtain hidden features corresponding to the time domain signal; The noise scene category of the original audio signal is predicted based on the hidden features through the output layer in the noise classification sub-model.
10. The method according to any one of claims 1 to 9, characterized in that: The method further comprises: Get the input noise reduction level; Determine a signal weight corresponding to the input noise reduction level, the signal weight comprising a first weight and a second weight respectively used to adjust a ratio between the original audio signal and the noise reduction signal; The original audio signal and the noise reduction signal are fused according to the first weight and the second weight to obtain an audio output signal corresponding to the input noise reduction level.
11. A method for processing an audio noise reduction model, characterized in that: The method comprises: Acquire a sample audio signal, where the sample audio signal is generated based on a clean audio signal; The real part processing network and the imaginary part processing network in the first sub-model based on the neural network are used to respectively perform feature encoding on the real part sequence and the imaginary part sequence obtained after the sample audio signal is transformed into a frequency domain signal, so as to obtain the real part attention and the imaginary part attention corresponding to the sample audio signal, and based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, obtain the frequency domain coding feature corresponding to the sample audio signal; the real part attention is used to reflect the attention to the clean signal in the real part frequency domain feature of the sample audio signal, and the imaginary part attention is used to reflect the attention to the clean signal in the imaginary part frequency domain feature of the sample audio signal; Performing model training on the first sub-model according to a first loss determined based on the frequency domain coding feature and a frequency domain transformation sequence corresponding to the clean audio signal to obtain a frequency domain processing sub-model; The frequency domain processing sub-model is connected to the time domain processing sub-model to be trained and trained together to obtain the audio noise reduction model for performing noise reduction processing on the audio signal.
12. The method according to claim 11, characterized in that The frequency domain processing sub-model is connected with the time domain processing sub-model to be trained and then trained together to obtain the audio noise reduction model for performing noise reduction processing on the audio signal, including: Connecting the frequency domain processing sub-model with the second sub-model and the third sub-model based on the neural network; Transforming the frequency domain coding features corresponding to the sample audio signal into a time domain signal; Constructing a multi-task objective function based on a second loss determined by a denoised signal obtained by performing denoising on the time domain signal by the second sub-model and the clean audio signal, and a third loss determined by a noise scene category obtained by performing noise classification on the time domain signal by the third sub-model and the noise label category of the sample audio signal; Performing model training on the frequency domain processing sub-model, the second sub-model and the third sub-model together according to the multi-task objective function to obtain an updated frequency domain processing sub-model, a trained time domain processing sub-model and a trained noise classification sub-model; The updated frequency domain processing sub-model is connected to the trained time domain processing sub-model to obtain the audio noise reduction model for performing noise reduction processing on the audio signal.
13. The method according to claim 12, characterized in that The steps of constructing the multi-task objective function include: Acquire multiple noise reduction signals obtained by performing noise reduction processing on time domain signals corresponding to multiple sample audio signals by the second sub-model, and determine the weight of the second loss according to standard deviations of the multiple noise reduction signals; Acquire multiple noise scene categories obtained by performing noise classification on time domain signals corresponding to multiple sample audio signals by the third sub-model, and determine the weight of the third loss according to the standard deviation of the multiple noise scene categories; The second loss and the third loss are fused according to their respective weights to obtain the multi-task objective function.
14. The method according to claim 11, characterized in that The frequency domain processing sub-model is connected with the time domain processing sub-model to be trained and then trained together to obtain the audio noise reduction model for performing noise reduction processing on the audio signal, including: Connecting the frequency domain processing sub-model with the second sub-model and the third sub-model based on the neural network; Transforming the frequency domain coding features corresponding to the sample audio signal into a time domain signal; Based on the noise reduction signal obtained by performing noise reduction processing on the time domain signal by the second sub-model and the second loss determined by the clean audio signal, the frequency domain processing sub-model and the second sub-model are jointly trained to obtain an updated frequency domain processing sub-model and a trained time domain processing sub-model; Input the sample audio signal into the updated frequency domain processing sub-model to obtain corresponding frequency domain coding features, transform the frequency domain coding features into time domain signals and input them into the trained time domain processing sub-model and the third sub-model respectively; Constructing a multi-task objective function based on a second loss determined by a denoised signal obtained by performing denoising processing on the time domain signal by the trained time domain processing sub-model and the clean audio signal, and a third loss determined by a noise scene category obtained by performing noise classification on the time domain signal by the third sub-model and the noise label category of the sample audio signal; Performing model training on the frequency domain processing sub-model, the trained time domain processing sub-model and the third sub-model according to the multi-task objective function to obtain an updated frequency domain processing sub-model, an updated time domain processing sub-model and a noise classification sub-model; The updated frequency domain processing sub-model is connected to the updated time domain processing sub-model to obtain the audio noise reduction model used for performing noise reduction processing on the audio signal.
15. The method according to any one of claims 11 to 14, characterized in that The method further comprises: A teacher model for performing noise reduction processing on an audio signal is obtained according to the audio noise reduction model, and a student model for performing noise reduction processing on an audio signal is constructed according to the frequency domain processing sub-model and the lightweight time domain noise reduction network; Input the sample audio signal into the teacher model, obtain the corresponding frequency domain coding features through the frequency domain processing sub-model in the teacher model, transform the frequency domain coding features into a time domain signal, encode the time domain signal through the encoder in the time domain processing sub-model in the teacher model to obtain a first time domain coding vector corresponding to the sample audio signal, and obtain a first noise reduction signal corresponding to the sample audio signal based on the first time domain coding vector through the decoder in the time domain processing sub-model; Input a sample audio signal into the student model, obtain a corresponding frequency domain coding feature through the frequency domain processing sub-model in the student model, transform the frequency domain coding feature into a time domain signal, encode the time domain signal through the encoder in the lightweight time domain denoising network, obtain a second time domain coding vector corresponding to the sample audio signal, and obtain a second denoised signal corresponding to the sample audio signal based on the second time domain coding vector through the decoder in the lightweight time domain denoising network; Based on the model distillation loss determined based on the mean square error loss between the first time domain coding vector and the second time domain coding vector, the mean square error loss between the first denoised signal and the second denoised signal, and the data structure loss between the first denoised signal and the second denoised signal, the student model is trained according to the model distillation loss to obtain a lightweight audio denoising model.
16. An audio noise reduction device, characterized in that: The device comprises: An acquisition module, used for acquiring the original audio signal in the time domain; A frequency domain coding module, used to perform feature coding on the real part sequence and the imaginary part sequence obtained after transforming the original audio signal into the frequency domain signal through the real part processing network and the imaginary part processing network in the trained frequency domain processing sub-model, respectively, to obtain the real part attention and the imaginary part attention corresponding to the original audio signal, wherein the real part attention is used to reflect the attention to the clean signal in the real part frequency domain features of the original audio signal, and the imaginary part attention is used to reflect the attention to the clean signal in the imaginary part frequency domain features of the original audio signal; based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention, the frequency domain coding feature corresponding to the original audio signal is obtained; The time domain noise reduction module is used to transform the frequency domain coding features into a time domain signal, and perform noise reduction processing on the time domain signal through a trained time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal.
17. The device according to claim 16, characterized in that The frequency domain encoding module is also used to input the original audio signal into a trained frequency domain processing sub-model; in the frequency domain processing sub-model, the original audio signal is transformed in the frequency domain to obtain a real sequence and an imaginary sequence corresponding to the original audio signal; through the real processing network in the frequency domain processing sub-model, the real sequence and the imaginary sequence are feature encoded respectively to obtain a real first coding feature and an imaginary first coding feature; through the imaginary processing network in the frequency domain processing sub-model, the real sequence and the imaginary sequence are feature encoded respectively to obtain a real second coding feature and an imaginary second coding feature; according to the real first coding feature and the imaginary second coding feature, the real part attention corresponding to the original audio signal is obtained, and according to the real second coding feature and the imaginary first coding feature, the imaginary part attention corresponding to the original audio signal is obtained.
18. The device according to claim 16, characterized in that The real part processing network and the imaginary part processing network in the frequency domain processing sub-model include at least two layers; The frequency domain coding module is also used to perform feature encoding on the real part sequence and the imaginary part sequence respectively through the real part processing network of the first layer in the frequency domain processing submodel to obtain a real first coding feature and an imaginary first coding feature; perform feature encoding on the real part sequence and the imaginary part sequence respectively through the imaginary part processing network of the first layer in the frequency domain processing submodel to obtain a real second coding feature and an imaginary second coding feature; obtain the real part attention corresponding to the original audio signal at the first layer according to the real first coding feature and the imaginary second coding feature, and obtain the imaginary part attention corresponding to the original audio signal at the first layer according to the real second coding feature and the imaginary first coding feature; iteratively perform feature encoding on the real part attention and imaginary part attention corresponding to the previous layer respectively through the real part processing network and the imaginary part processing network of the current layer to obtain the real part attention and imaginary part attention corresponding to the current layer, and stop the iteration until the real part attention and imaginary part attention corresponding to the last layer are obtained.
19. The device according to claim 16, characterized in that The frequency domain coding module is also used to multiply the real part sequence by the real part attention to obtain the real part of the frequency domain coding feature corresponding to the original audio signal; and multiply the imaginary part sequence by the imaginary part attention to obtain the imaginary part of the frequency domain coding feature corresponding to the original audio signal.
20. The device according to claim 16, characterized in that The frequency domain coding module is also used to multiply the real part sequence by the real part attention to obtain a first result, multiply the imaginary part sequence by the imaginary part attention to obtain a second result, and use the difference between the first result and the second result as the real part of the frequency domain coding feature corresponding to the original audio signal; multiply the real part sequence by the imaginary part attention to obtain a third result, multiply the imaginary part sequence by the real part attention to obtain a fourth result, and use the sum of the third result and the fourth result as the imaginary part of the frequency domain coding feature corresponding to the original audio signal.
21. The device according to claim 16, characterized in that The frequency domain coding module is also used to obtain the original amplitude information and the original phase information of the original audio signal based on the real part sequence and the imaginary part sequence, and to obtain the predicted amplitude information and the predicted phase information of the original audio signal based on the real part attention and the imaginary part attention; to obtain the amplitude information of the frequency domain coding feature corresponding to the original audio signal according to the product of the original amplitude information and the predicted amplitude information; and to obtain the phase information of the frequency domain coding feature corresponding to the original audio signal according to the sum of the original phase information and the predicted phase information.
22. The device according to claim 16, characterized in that The time domain noise reduction module is also used to input the time domain signal into a trained time domain processing sub-model; encode the time domain signal through the encoder in the time domain processing sub-model to obtain a time domain coding vector; extract features from the time domain coding vector through the time series feature extraction network in the time domain processing sub-model to obtain hidden features corresponding to the time domain signal; and decode based on the time domain coding vector and the hidden features through the decoder in the time domain processing sub-model to obtain a noise reduction signal corresponding to the original audio signal.
23. The device according to claim 16, characterized in that The device also includes: The noise classification module is used to classify the time domain signal through a trained noise classification sub-model to obtain the noise scene category of the original audio signal. The noise classification sub-model is obtained by jointly training the model with the time domain processing sub-model.
24. The device according to claim 23, characterized in that The noise classification module is also used to input the time domain signal into a trained noise classification sub-model; encode the time domain signal through the encoder in the noise classification sub-model to obtain a time domain coding vector; extract features from the time domain coding vector through the temporal feature extraction network in the noise classification sub-model to obtain hidden features corresponding to the time domain signal; and predict the noise scene category of the original audio signal based on the hidden features through the output layer in the noise classification sub-model.
25. The device according to any one of claims 16 to 24, characterized in that The device also includes: A noise reduction gear adjustment module is used to obtain an input noise reduction level; determine a signal weight corresponding to the input noise reduction level, the signal weight including a first weight and a second weight respectively used to adjust the ratio between the original audio signal and the noise reduction signal; according to the first weight and the second weight, after fusing the original audio signal with the noise reduction signal, obtain an audio output signal corresponding to the input noise reduction level.
26. A processing device for an audio noise reduction model, characterized in that: The device comprises: An acquisition module, used for acquiring a sample audio signal, wherein the sample audio signal is generated according to a clean audio signal; A frequency domain coding training module is used to perform feature coding on the real part sequence and the imaginary part sequence obtained after transforming the sample audio signal into a frequency domain signal through the real part processing network and the imaginary part processing network in the first sub-model based on the neural network, respectively, to obtain the real part attention and the imaginary part attention corresponding to the sample audio signal, and obtain the frequency domain coding feature corresponding to the sample audio signal based on the real part sequence and the imaginary part sequence, the real part attention and the imaginary part attention; the first sub-model is model trained according to the first loss determined based on the frequency domain coding feature and the frequency domain transformation sequence corresponding to the clean audio signal to obtain the frequency domain processing sub-model; the real part attention is used to reflect the attention to the clean signal in the real part frequency domain feature of the sample audio signal, and the imaginary part attention is used to reflect the attention to the clean signal in the imaginary part frequency domain feature of the sample audio signal; The integrated training module is used to connect the frequency domain processing sub-model and the time domain processing sub-model to be trained and train them together to obtain the audio noise reduction model used for noise reduction processing of audio signals.
27. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method according to any one of claims 1 to 15 when executing the computer program.
28. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15.
29. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.
Citation Information
Patent Citations
Complex field speech enhancement method and system based on generative adversarial network and medium
CN110739002A
Voice processing method, voice processing device and device for processing voice
CN110808063A
Cited By
Speech enhancement method and system based on selective state-space model time-frequency multi-directional scanning
CN122799878A