Voice Processing Method, Apparatus, Device, and Storage Medium

By obtaining the frequency domain information of speech data for feature extraction and mask generation, the problem of deep learning algorithms consume high computing resources in speech enhancement is solved, and efficient speech noise reduction is achieved.

CN113823313BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110783691.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-12
Publication Date
2025-07-11
Estimated Expiration
2041-07-12

AI Technical Summary

Technical Problem

When the prior art uses deep learning algorithms for speech enhancement, computing resources consume a lot, resulting in high demand for computing resources.

Method used

By obtaining the frequency domain information of speech data, feature extraction and mask generation are performed, and noise is directly removed based on the frequency domain information, reducing dependence on complex models, improving speech noise reduction speed and reducing computing resource consumption.

Benefits of technology

While ensuring the noise reduction effect, the speed of voice noise reduction is improved and the consumption of computing resources is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113823313B_ABST
    Figure CN113823313B_ABST
Patent Text Reader

Abstract

The present application discloses a voice processing method, apparatus, device, and storage medium, belonging to the field of computer technology. Through the technical solution provided by the embodiments of the present application, when performing voice noise reduction, there is no need to identify noise through a model with a complex structure. Instead, a first mask is directly determined based on the frequency domain information of the voice data, and by combining the first mask with the spectrum of the voice data, the target voice data can be obtained, which improves the speed of voice noise reduction and reduces the consumption of computing resources while ensuring the noise reduction effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a voice processing method, apparatus, device, and storage medium. Background Art

[0002] Voice enhancement refers to extracting useful original voice from the noise background when voice data is interfered with or even submerged by various noises, so as to suppress and reduce noise interference. In short, voice enhancement refers to extracting as pure original voice as possible from noisy voice.

[0003] In related technologies, voice enhancement is performed based on deep learning algorithms. For example, voice activity detection (VAD) is performed on voice data through deep learning algorithms, and then noise is identified from the voice data based on the VAD results, and the voice data is subtracted from the noise to obtain pure original voice.

[0004] However, when using deep learning algorithms for voice enhancement, due to multiple iterative processes involved, the consumption of computing resources is relatively large. Summary of the Invention

[0005] Embodiments of this application provide a voice processing method, apparatus, device, and storage medium, which can reduce the consumption of computing resources while ensuring the voice noise reduction effect. The technical solution is as follows:

[0006] On the one hand, a voice processing method is provided, and the method includes:

[0007] Obtain multiple frequency domain information of voice data, and the multiple frequency domain information corresponds to multiple audio frames of the voice data one by one;

[0008] Extract features from the multiple frequency domain information to obtain multiple first frequency domain features, and the first frequency domain features are determined based on the corresponding audio frames and the audio frames adjacent to the audio frames;

[0009] Based on the multiple first frequency domain features, obtain a first mask of the voice data, and the first mask is used to remove noise from the voice data;

[0010] Generate target voice data based on the spectrum of the voice data and the first mask.

[0011] On the one hand, a voice processing apparatus is provided, and the apparatus includes:

[0012] A frequency domain information acquisition module, configured to obtain multiple frequency domain information of voice data, and the multiple frequency domain information corresponds to multiple audio frames of the voice data one by one;

[0013] A feature extraction module, configured to extract features from the multiple pieces of frequency-domain information to obtain multiple first frequency-domain features, where the first frequency-domain features are determined based on corresponding audio frames and audio frames adjacent to the audio frames;

[0014] A first mask acquisition module, configured to obtain a first mask of the speech data based on the multiple first frequency-domain features, where the first mask is used to remove noise from the speech data;

[0015] A speech data generation module, configured to generate target speech data based on the spectrum of the speech data and the first mask.

[0016] In a possible implementation manner, the frequency-domain information acquisition module is configured to obtain the mean and variance of the multiple pieces of initial frequency-domain information; and use the mean and the variance to perform normalization processing on the multiple pieces of initial frequency-domain information to obtain the multiple pieces of frequency-domain information.

[0017] In a possible implementation manner, the feature extraction module is configured to input the multiple pieces of frequency-domain information into a speech enhancement model, and through the speech enhancement model, extract features from the multiple pieces of frequency-domain information to obtain multiple second frequency-domain features; and through the speech enhancement model, based on the multiple second frequency-domain features, obtain the multiple first frequency-domain features in the arrangement order of the multiple audio frames.

[0018] In a possible implementation manner, for any one of the multiple audio frames, the feature extraction module is configured to obtain the first frequency-domain feature of the audio frame based on the second frequency-domain feature of the audio frame and the second frequency-domain features of at least one audio frame adjacent to the audio frame.

[0019] In a possible implementation manner, the first mask acquisition module is configured to perform a fully connected process on the multiple first frequency-domain features through a speech enhancement model to obtain the first mask of the speech data.

[0020] In a possible implementation manner, the speech data generation module is configured to multiply multiple frequency points of the spectrum of the speech data by the first mask through a speech enhancement model to obtain a first target spectrum; and convert the first target spectrum into the target speech data.

[0021] In a possible implementation manner, the apparatus further includes:

[0022] A model training module, configured to obtain first sample voice data and second sample voice data, where the second sample voice data is voice data obtained by adding noise to the first sample voice data; input the second sample voice data into the voice enhancement model, and through the voice enhancement model, obtain a predicted first mask of the second sample voice data; multiply multiple frequency points of the spectrum of the second sample voice data by the predicted first mask to obtain a predicted spectrum; train the voice enhancement model based on the difference information between the predicted spectrum and the spectrum of the first sample voice data.

[0023] In a possible implementation manner, the apparatus further includes:

[0024] A second mask obtaining module, configured to perform statistical noise reduction on voice data to obtain a second mask of the voice data;

[0025] The voice data generation module is further configured to generate the target voice data based on the second mask, the spectrum of the labeled voice data, and the first mask.

[0026] In a possible implementation manner, the second mask obtaining module is configured to perform any one of the following:

[0027] Obtain a noise estimation spectrum of the voice data; based on the noise estimation spectrum and the spectrum of the voice data, obtain the second mask;

[0028] Perform Wiener filtering on the voice data to obtain the second mask.

[0029] In a possible implementation manner, the voice data generation module is further configured to fuse the first mask and the second mask to obtain a fused mask; multiply multiple frequency points of the spectrum of the target voice data by the fused mask to obtain a second target spectrum; convert the second spectrum into the target voice data.

[0030] On the one hand, a computer device is provided, where the computer device includes one or more processors and one or more memories, and at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the voice processing method.

[0031] On the one hand, a computer-readable storage medium is provided, where at least one computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the voice processing method.

[0032] On the one hand, a computer program product or a computer program is provided. The computer program product or the computer program includes program code stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the above-mentioned voice processing method.

[0033] Through the technical solution provided by the embodiments of the present application, when performing voice noise reduction, it is not necessary to identify noise through a model with a complex structure. A first mask is directly determined based on the frequency domain information of the voice data, and the first mask is combined with the spectrum of the voice data to obtain the target voice data. While ensuring the noise reduction effect, the speed of voice noise reduction is improved, and the consumption of computing resources is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0035] Figure 1 is a schematic diagram of the implementation environment of a voice processing method provided by the embodiments of the present application;

[0036] Figure 2 is a flowchart of a voice processing method provided by the embodiments of the present application;

[0037] Figure 3 is a flowchart of a voice processing method provided by the embodiments of the present application;

[0038] Figure 4 is a schematic structural diagram of a voice enhancement model provided by the embodiments of the present application;

[0039] Figure 5 is a flowchart of a voice processing method provided by the embodiments of the present application;

[0040] Figure 6 is a flowchart of a voice processing method provided by the embodiments of the present application;

[0041] Figure 7 is a flowchart of a training method for a voice enhancement model provided by the embodiments of the present application;

[0042] Figure 8 is a flowchart of a voice processing method provided by the embodiments of the present application;

[0043] Figure 9It is a flowchart of a voice processing method provided by an embodiment of the present application;

[0044] Figure 10 It is a schematic structural diagram of a voice processing device provided by an embodiment of the present application;

[0045] Figure 11 It is a schematic structural diagram of a terminal provided by an embodiment of the present application;

[0046] Figure 12 It is a schematic structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners

[0047] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0048] In the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions. It should be understood that there is no logical or chronological dependency between "first", "second", and "nth", nor are the quantity and execution order limited.

[0049] In the present application, the term "at least one" means one or more, and the meaning of "multiple" means two or more. For example, multiple face images mean two or more face images.

[0050] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.

[0051] Cloud Computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to users to be infinitely scalable, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. As a basic capability provider of cloud computing, a cloud computing resource pool (abbreviated as a cloud platform, generally called an IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (virtual machines, including operating systems), storage devices, and network devices. According to logical function division, a Platform as a Service (PaaS) layer can be deployed on the Infrastructure as a Service (IaaS) layer, and a Software as a Service (SaaS) layer can be deployed on top of the PaaS layer. It is also possible to directly deploy SaaS on IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is various business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.

[0052] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0053] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0054] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize existing knowledge sub-models to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0055] Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.

[0056] Figure 1 It is a schematic diagram of the implementation environment of a voice processing method provided by an embodiment of this application. Refer to Figure 1 In this implementation environment, it may include a first terminal 110, a second terminal 120, and a server 140.

[0057] Optionally, the first terminal 110 is a smartphone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. An application program that supports voice calls is installed and run on the first terminal 110.

[0058] Optionally, the second terminal 120 is a smartphone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. An application program that supports voice calls is installed and run on the second terminal 120.

[0059] Optionally, the first terminal 110 can be connected to the server 140 through a wireless network or a wired network. The first terminal 110 can send the collected voice message to the server 140, and the server 140 sends the voice message to the second terminal 120. The user of the second terminal 120 can listen to the voice message through the second terminal 120. If the user of the second terminal 120 is inconvenient to listen to the voice message, then the user of the second terminal 120 can send the voice message to the server through the second terminal 120. The server performs voice enhancement on the voice message and sends the result of the voice enhancement to the second terminal 120, and the second terminal 120 presents the result of the voice enhancement to the user.

[0060] Optionally, the server 140 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.

[0061] Optionally, the above-mentioned first terminal 110, second terminal 120, and server 140 can act as nodes on the blockchain system.

[0062] It should be noted that the voice processing method provided in the embodiments of the present application can be executed by the terminal or by the server. The embodiments of the present application do not make any limitations in this regard. When the voice processing method provided in the embodiments of the present application is executed by the server, the voice processing method can be used as a cloud service, and the terminal can use the cloud service through the server.

[0063] After introducing the implementation environment of the embodiments of the present application, the application scenarios of the embodiments of the present application will be introduced below in combination with the above implementation environment. It should be noted that in the following description process, the terminal is the above-mentioned first terminal 110 or second terminal 120, and the server is the above-mentioned server 140.

[0064] The voice processing method provided in the embodiments of the present application can be applied to the scenario of real-time voice calls. That is, through the voice processing method provided in the embodiments of the present application, when two or more parties are having an online voice call, the voice data of each party can be enhanced to reduce the noise in the voice data. For example, it can be applied to the scenario of making a call through an instant messaging application, or the scenario of making a call through an online meeting application, or the scenario of making a voice call through a phone based on VOIP (Voice over Internet Protocol). In addition, the voice processing method provided in the embodiments of the present application can also be applied to other scenarios that require noise reduction, such as the scenario of headphone noise reduction or hearing aid noise reduction, or the scenario of noise reduction for audio files. Of course, with the development of science and technology, the voice processing method provided in the embodiments of the present application can also be applied to other scenarios that require voice enhancement. The embodiments of the present application do not make any limitations in this regard.

[0065] In the scenario of making a call through an instant messaging application, taking the call between two users as an example, the two users respectively use the first terminal 110 and the second terminal 120 in the above implementation environment to make a voice call. For the user using the first terminal 110, the user starts the instant messaging application installed on the first terminal 110, and selects the user with whom they want to make a voice call through this instant messaging application. This user is also the user using the second terminal 120, and this instant messaging application has the function of voice call. During the call between the two users, the first terminal 110 can collect the voice signals emitted by the user, execute the voice processing method provided by the embodiments of the present application on this voice data, perform voice enhancement on the collected voice data, and obtain the first target voice data. The first terminal 110 sends the first target voice data to the server 140, and the server 140 sends the first target voice data to the second terminal 120. The second terminal 120 plays the first target voice data, and the user using the second terminal 120 can thus hear the content expressed by the user using the first terminal 110. Correspondingly, during the call between the two users, the user using the second terminal 120 will also express content to the user using the first terminal 110. In this case, the second terminal 120 can collect the voice signals emitted by the user, execute the voice processing method provided by the embodiments of the present application on this voice data, perform voice enhancement on the collected voice data, and obtain the first target voice data. The second terminal 120 sends the first target voice data to the server 140, and the server 140 sends the first target voice data to the first terminal 110. The first terminal 110 plays the first target voice data, and the user using the first terminal 110 can thus hear the content expressed by the user using the second terminal 120.

[0066] It should be noted that the above is described by taking the first terminal 110 and the second terminal 120 as examples for executing the voice processing method provided by the embodiments of the present application. In other possible implementation manners, the voice processing method provided by the embodiments of the present application can also be executed by the server 140. That is, when two users make a voice call through the first terminal 110 and the second terminal 120, the first terminal 110 and the second terminal 120 directly send the collected voice data to the server 140. The server 140 executes the voice processing method provided by the embodiments of the present application on the received voice data, reduces the noise in the voice data, and sends the voice data after voice enhancement to the first terminal 110 or the second terminal 120.

[0067] In the scenario of making a call through an online meeting application, there are often multiple participants in the online meeting. In some embodiments, the first terminal 110 is the initiator of the online meeting, and multiple second terminals 120 are the participants in the online meeting. For the user using the first terminal 110, they can start the online meeting application through the first terminal 110 and create an online meeting through the online meeting application. The user using the first terminal 110 can share the link of the online meeting with multiple second terminals 120, and the second terminals 120 can join the online meeting based on the link of the online meeting. During the online meeting, both the first terminal 110 and multiple second terminals 120 are participants in the online meeting. The first terminal 110 can collect the voice signal emitted by the user, execute the voice processing method provided by the embodiments of the present application on the voice data, perform voice enhancement on the collected voice data, and obtain the first target voice data. The first terminal 110 sends the first target voice data to the server 140, and the server 140 sends the first target voice data to the second terminal 120. The second terminal 120 plays the first target voice data, and the user using the second terminal 120 can also hear the content expressed by the user using the first terminal 110. Correspondingly, during the call between two users, the user using the second terminal 120 will also express content to the user using the first terminal 110. In this case, the second terminal 120 can collect the voice signal emitted by the user, execute the voice processing method provided by the embodiments of the present application on the voice data, perform voice enhancement on the collected voice data, and obtain the first target voice data. The second terminal 120 sends the first target voice data to the server 140, and the server 140 sends the first target voice data to the first terminal 110. The first terminal 110 plays the first target voice data, and the user using the first terminal 110 can also hear the content expressed by the user using the second terminal 120. Of course, since there may be many participants in the online meeting, the terminals used by each participant can all adopt the voice processing method provided by the embodiments of the present application to perform voice enhancement on the voice data.

[0068] In the scenario of hearing aid noise reduction, when the user wearing the hearing aid turns on the voice noise reduction function of the hearing aid, the hearing aid can adopt the voice processing method provided by the embodiments of the present application to perform voice enhancement on the collected voice data to reduce the noise in the voice data.

[0069] In the scenario of performing noise reduction on an audio file, the terminal can obtain the audio file, execute the voice processing method provided by the embodiments of the present application on the audio signal stored in the audio file, perform voice enhancement on the audio signal, and generate a target audio file based on the voice-enhanced audio signal. The target audio file is also the audio file after voice enhancement.

[0070] After introducing the implementation environment and application scenarios of the embodiments of the present application, the voice processing method provided by the embodiments of the present application will be described below.

[0071] Figure 2 It is a flowchart of a voice processing method provided by an embodiment of the present application. Taking the execution entity as a terminal as an example, see Figure 2 The method includes:

[0072] 201. The terminal obtains multiple frequency-domain information of the voice data, and the multiple frequency-domain information corresponds one-to-one to multiple audio frames of the voice data.

[0073] Among them, the frequency-domain information is also the representation of the audio frame in the frequency domain. The one-to-one correspondence between the multiple frequency-domain information and the multiple audio frames means that each audio frame corresponds to a frequency-domain information.

[0074] 202. The terminal extracts features from the multiple frequency-domain information to obtain multiple first frequency-domain features, and the first frequency-domain features are determined based on the corresponding audio frames and the audio frames adjacent to the audio frames.

[0075] Among them, the first frequency-domain features are determined based on the corresponding audio frames and the audio frames adjacent to the audio frames, which means that the first frequency-domain features are a fusion feature, or the first frequency-domain features are a feature that combines "context information". For the i-th audio frame, the above context information is the (i - 1)-th audio frame, and the below context information is the (i + 1)-th audio frame, where i is the serial number of the audio frame and i is a positive integer.

[0076] 203. The terminal obtains a first mask of the voice data based on the multiple first frequency-domain features, and the first mask is used to remove noise in the voice data.

[0077] In some embodiments, the first mask is a matrix, and the values in the matrix are all within the range of 0 to 1. The value 0 indicates that the corresponding frequency point is a noise frequency point and needs to be completely eliminated, the value 1 indicates that the corresponding frequency point is a non-noise frequency point and does not need to be eliminated, and the value between 0 and 1 indicates that the corresponding frequency point is a frequency point containing noise and needs to be partially eliminated. The noise in the voice data can be eliminated through the first mask.

[0078] 204. The terminal generates target voice data based on the spectrum of the voice data and the first mask.

[0079] Among them, the target voice data is also the voice data obtained after voice enhancement of the voice data.

[0080] Through the technical solution provided by the embodiments of the present application, when performing speech noise reduction, there is no need to identify noise through a model with a complex structure. Instead, a first mask is directly determined based on the frequency-domain information of the speech data, and by combining the first mask with the spectrum of the speech data, the target speech data can be obtained, which improves the speed of speech noise reduction and reduces the consumption of computing resources while ensuring the noise reduction effect.

[0081] The above steps 201-204 are a simple description of the speech processing method provided by the embodiments of the present application. Next, some examples will be combined to provide a detailed description of the speech processing method provided by the embodiments of the present application.

[0082] Figure 3 is a flowchart of a speech processing method provided by the embodiments of the present application. Refer to Figure 3 , taking the first terminal sending speech data to the second terminal as an example, the method includes:

[0083] 301. The first terminal acquires speech data.

[0084] In a possible implementation manner, the first terminal collects the sound emitted by the user through a microphone and converts the sound emitted by the user into speech data. This implementation manner can be applied in scenarios such as real-time voice calls, headphone noise reduction, or hearing aid noise reduction. In these scenarios, the first terminal executes the speech processing method provided by the embodiments of the present application for the real-time collected audio signal. Therefore, the first terminal can collect speech data in real time through the microphone.

[0085] In a possible implementation manner, the first terminal loads an audio file and acquires speech data from the audio file. In this implementation manner, the first terminal can execute the speech processing method provided by the embodiments of the present application for the audio signal stored in the audio file. For example, when performing post-processing on some noisy speech files, speech data can be acquired through this implementation manner.

[0086] 302. The first terminal acquires multiple frequency-domain information of the speech data, and the multiple frequency-domain information corresponds to multiple audio frames of the speech data one by one.

[0087] Among them, the frequency-domain information is also the representation of the audio frame in the frequency domain.

[0088] In a possible implementation manner, the first terminal frames and windows the speech data to obtain multiple audio frames. The first terminal performs time-frequency transformation on the multiple audio frames to obtain multiple initial frequency-domain information respectively corresponding to the multiple audio frames. The first terminal normalizes the multiple initial frequency-domain information to obtain the multiple frequency-domain information.

[0089] For a clearer description of the above embodiments, the above embodiments will be described in three parts below.

[0090] In the first part, the first terminal frames and windows the voice data to obtain a plurality of audio frames.

[0091] In a possible implementation manner, the first terminal divides the voice data into a plurality of initial audio frames based on a target frame length and a target frame shift. The first terminal performs windowing processing on each initial audio frame to obtain a plurality of audio frames. Among them, the target frame length is set according to the actual situation. For example, the range of the target frame length is 5-30 ms. The target frame shift can be set according to the target frame length. For example, it is one-half or one-third of the target frame length. The present application embodiment does not limit the setting method and size of the target frame length and the target frame shift. In some embodiments, after framing, there is an overlapping part between two adjacent audio frames, which can ensure the integrity of the information recorded in the audio frames.

[0092] Since the audio signal changes greatly over time, overall, the audio signal is unstable. For the convenience of analysis, the computer device can frame the received audio signal, decompose the overall unstable audio signal into a plurality of audio frames, and the plurality of audio frames can be considered stable locally. After the segmentation points of the plurality of audio frames are subjected to time-frequency conversion, they will present high-frequency components in the frequency domain, and these high-frequency components do not exist in the original audio signal. To reduce the influence of this situation, by performing windowing processing on each audio frame, the separation points between two adjacent audio frames become continuous.

[0093] For example, the first terminal divides the voice data into a plurality of initial voice frames based on a target frame length and a target frame shift. The target frame shift refers to the interval between two adjacent audio frames. For example, the target frame length is 8 ms. If the overlap degree between two adjacent audio frames is 80%, it means that the first terminal obtains an audio frame with a frame length of 8 ms every 2 ms. Among them, 2 ms is also the target frame shift. The first audio frame is the part of the voice data from 0 ms to 8 ms, the second audio frame is the part of the voice data from 2 ms to 10 ms, the third audio frame is the part of the voice data from 4 ms to 12 ms, and so on. The first terminal uses a window function to perform windowing processing on each initial voice frame, that is, multiplies the window function by each initial voice frame to obtain a plurality of audio frames. In some embodiments, the window function includes a Hamming window, a Hanning window, a triangular window, a Blackman window, and a Kaiser window, etc. The present application embodiment does not limit this.

[0094] In the second part, the first terminal performs time-frequency transformation on the plurality of audio frames to obtain a plurality of initial frequency domain information.

[0095] Among them, multiple initial frequency domain information corresponds to multiple audio frames one by one. In some embodiments, the initial frequency domain information is also referred to as logarithmic spectrum information. The fact that multiple initial frequency domain information and multiple audio frames correspond to each other one by one means that each audio frame corresponds to one initial frequency domain information.

[0096] In a possible implementation manner, the first terminal performs Fourier transform on the multiple audio frames to obtain multiple initial frequency domain information.

[0097] Next, taking the window function adopted by the first terminal as the Hanning window and the target frame length as 20 ms as an example, the method for obtaining the initial frequency domain information of the audio frame in the embodiments of the present application will be described.

[0098] The window function of the Hanning window is formula (1).

[0099]

[0100] The formula for Fourier transform is formula (2).

[0101]

[0102] Among them, win(n) is the window function of the Hanning window, X(i, k) is the initial frequency domain information, x(n) is the audio signal, n is the length of the audio frame, N is the number of audio frames, N is a positive integer, k is the serial number of the index point, k = 1, 2, 3... N, i is the serial number of the audio frame, and j represents a complex number.

[0103] It should be noted that in addition to being able to perform time-frequency transformation on multiple audio frames through the method of Fourier transform, the first terminal can also perform time-frequency transformation on multiple audio frames through short-time Fourier transform (STFT), wavelet algorithm or fast Fourier transform (FFT). The embodiments of the present application do not make limitations in this regard.

[0104] Part Three: The first terminal normalizes the multiple initial frequency domain information to obtain the multiple frequency domain information.

[0105] In a possible implementation manner, the first terminal obtains the mean and variance of the multiple initial frequency domain information. The first terminal uses the mean and the variance to normalize the multiple initial frequency domain information to obtain the multiple frequency domain information.

[0106] For example, the first terminal obtains the mean of multiple initial frequency domain information, and based on the multiple initial frequency domain information and the mean, obtains the variance of the multiple initial frequency domain information. The first terminal subtracts the multiple initial frequency domain information from the mean respectively to obtain a first difference. The first terminal divides the first difference by the standard deviation to obtain the multiple frequency domain information, where the standard deviation is obtained by taking the square root of the variance.

[0107] 303. The first terminal extracts features from the multiple frequency domain information to obtain multiple first frequency domain features, and the first frequency domain features are determined based on the corresponding audio frame and the audio frames adjacent to the audio frame.

[0108] In a possible implementation manner, the first terminal inputs the multiple frequency domain information into a voice enhancement model, and through the voice enhancement model, extracts features from the multiple frequency domain information to obtain multiple second frequency domain features. The first terminal, through the voice enhancement model, according to the arrangement order of the multiple audio frames, obtains the multiple first frequency domain features based on the multiple second frequency domain features.

[0109] To illustrate the above implementation manner more clearly, the above implementation manner will be described in two parts below.

[0110] The first part: The first terminal inputs the multiple frequency domain information into a voice enhancement model, and through the voice enhancement model, extracts features from the multiple frequency domain information to obtain multiple second frequency domain features.

[0111] In a possible implementation manner, the first terminal inputs multiple frequency domain information into a voice enhancement model, and through the feature extraction layer of the voice enhancement model, sequentially extracts features from the multiple frequency domain information, and combines the result of the previous feature extraction each time when extracting features, to obtain multiple second frequency domain features. In some embodiments, the voice enhancement model is a recurrent neural network (RNN), such as a long short-term memory network (LSTM), or a gated recurrent unit (GRU), etc., and the embodiments of the present application do not limit this.

[0112] For example, for the frequency domain information of the i-th audio frame, the first terminal inputs the frequency domain information of the i-th audio frame into the first feature extraction layer of the voice enhancement model, and the first feature extraction layer includes multiple computing units. In some embodiments, the i-th audio frame is the first audio frame, and the first terminal extracts features from the frequency domain information of the i-th audio frame through the first computing unit in the first feature extraction layer, that is, multiplies the frequency domain information of the i-th audio frame by the first weight matrix to obtain the second frequency domain feature of the frequency domain information of the i-th audio frame, and the first weight matrix is the weight matrix corresponding to the first feature extraction layer of the voice enhancement model. For the frequency domain information of the (i + 1)-th audio frame, the first terminal inputs the frequency domain information of the (i + 1)-th audio frame into the first feature extraction layer of the voice enhancement model, and extracts features from the frequency domain information of the (i + 1)-th audio frame through the second computing unit in the first feature extraction layer to obtain the second frequency domain feature of the frequency domain information of the (i + 1)-th audio frame. When the first terminal extracts features from the frequency domain information of the (i + 1)-th audio frame through the second computing unit, in addition to inputting the frequency domain information of the (i + 1)-th audio frame into the second computing unit, the first terminal also inputs the output of the first computing unit, that is, the second frequency domain feature of the frequency domain information of the i-th audio frame, into the second computing unit. The second computing unit fuses the second frequency domain feature of the first frequency domain information with the frequency domain information of the (i + 1)-th audio frame, multiplies the fused information by the first weight matrix to obtain the second frequency domain feature of the frequency domain information of the (i + 1)-th audio frame. For the frequency domain information of the (i + 2)-th audio frame, the first terminal inputs the frequency domain information of the (i + 2)-th audio frame into the first feature extraction layer of the voice enhancement model, and extracts features from the frequency domain information of the (i + 2)-th audio frame through the third computing unit in the first feature extraction layer to obtain the second frequency domain feature of the frequency domain information of the (i + 2)-th audio frame. When the first terminal extracts features from the frequency domain information of the (i + 2)-th audio frame through the third computing unit, in addition to inputting the frequency domain information of the (i + 2)-th audio frame into the third computing unit, the first terminal also inputs the output of the second computing unit, that is, the second frequency domain feature of the frequency domain information of the (i + 1)-th audio frame, into the third computing unit. The third computing unit fuses the second frequency domain feature of the frequency domain information of the (i + 1)-th audio frame with the frequency domain information of the (i + 2)-th audio frame, multiplies the fused information by the first weight matrix to obtain the second frequency domain feature of the frequency domain information of the (i + 2)-th audio frame. And so on, until multiple second frequency domain features corresponding to multiple frequency domain information are obtained.

[0113] Figure 4 A schematic structural diagram of a voice enhancement model is given, see Figure 4, the voice enhancement model 400 includes an input layer 401, a first feature extraction layer 402, a second feature extraction layer 403, and an output layer 404. Corresponding to the above embodiments, the first terminal inputs multiple frequency domain information into the voice enhancement model 400 through the input layer 401, and through the first feature extraction layer 402, sequentially extracts features from the multiple frequency domain information to obtain multiple second frequency domain features. Figure 4 W1 in

[0114] In some embodiments, when the subsequent computing unit of the first feature extraction layer extracts features from the frequency domain information, it combines the output of the previous computing unit. To ensure that the gradient of the model does not explode, the first terminal can perform a "forgetting" process on the output of the previous computing unit, that is, multiply the output of the previous computing unit by a weight, and the value range of this weight is 0 to 1. When the weight is 0, it means that the output of the previous computing unit is completely "forgotten". When the weight is 1, it means that the output of the previous computing unit is completely retained. The first terminal can input the result after the "forgetting" process into the subsequent computing unit to avoid gradient explosion.

[0115] In a possible implementation manner, the first terminal inputs multiple frequency domain information into the voice enhancement model, and through the feature extraction layer of the voice enhancement model, performs convolution processing on the multiple frequency domain information to obtain multiple second frequency domain features. In some embodiments, the voice enhancement model is a Convolutional Neural Networks (CNN).

[0116] For example, the first terminal inputs multiple frequency domain information into the voice enhancement model, and through the first convolutional kernel of the voice enhancement model, performs convolution processing on the multiple frequency domain information respectively to obtain multiple second frequency domain features, and the first convolutional kernel is the convolutional kernel corresponding to the first feature extraction layer of the voice enhancement model.

[0117] In a possible implementation manner, the first terminal inputs multiple frequency domain information into the voice enhancement model, and through the feature extraction layer of the voice enhancement model, encodes the multiple frequency domain information based on the attention mechanism to obtain multiple second frequency domain features.

[0118] For example, the first terminal inputs multiple frequency domain information into the voice enhancement model, and through the feature extraction layer of the voice enhancement model, obtains the key (K, Key) matrix, query (Q, Query) matrix, and value (V, Value) matrix of each frequency domain information. The first terminal determines the second frequency domain features of each frequency domain information based on the key matrix, query matrix, and value matrix of each frequency domain information.

[0119] In the second part, the first terminal uses the voice enhancement model to obtain the plurality of first frequency domain features based on the plurality of second frequency domain features in the arranged order of the plurality of audio frames.

[0120] In a possible implementation manner, for any one of the plurality of audio frames, the first terminal obtains the first frequency domain feature of the audio frame based on the second frequency domain feature of the audio frame and the second frequency domain features of at least one audio frame adjacent to the audio frame.

[0121] For example, the first terminal inputs the plurality of second frequency domain features into the second feature extraction layer of the voice enhancement model. The second feature extraction layer obtains the first frequency domain feature corresponding to each second frequency domain feature based on each second frequency domain feature and the second frequency domain features of at least one reference audio frame. Each second frequency domain feature corresponds to a target audio frame, the reference audio frame is the audio frame adjacent to the target audio frame, and the first frequency domain feature corresponding to each second frequency domain feature is also the first frequency domain feature corresponding to the target audio frame.

[0122] For example, for the second frequency domain feature of the i-th audio frame, the first terminal inputs the second frequency domain feature of the i-th audio frame into the second feature extraction layer, and the second feature extraction layer includes multiple computing units. The first terminal extracts the feature of the second frequency domain feature of the i-th audio frame through the first computing unit of the second feature extraction layer to obtain the first frequency domain feature of the i-th audio frame. When the first terminal extracts the feature of the second frequency domain feature of the i-th audio frame through the first computing unit in the second feature extraction layer, in addition to inputting the second frequency domain feature of the i-th audio frame into the first computing unit, the first terminal also inputs the second frequency domain feature of the (i + 1)-th audio frame into the first computing unit. The first computing unit fuses the second frequency domain feature of the i-th audio frame and the second frequency domain feature of the (i + 1)-th audio frame, multiplies the fused information by the second weight matrix to obtain the first frequency domain feature of the (i + 1)-th audio frame, and the second weight matrix is the weight matrix corresponding to the second feature extraction layer of the speech enhancement model. For the second frequency domain feature of the (i + 1)-th audio frame, the first terminal inputs the second frequency domain feature of the (i + 1)-th audio frame into the second feature extraction layer of the speech enhancement model, and extracts the feature of the second frequency domain feature of the (i + 1)-th audio frame through the second computing unit in the second feature extraction layer to obtain the first frequency domain feature of the (i + 1)-th audio frame. When the first terminal extracts the feature of the second frequency domain feature of the (i + 1)-th audio frame through the second computing unit, in addition to inputting the second frequency domain feature of the (i + 1)-th audio frame into the second computing unit, the first terminal also inputs the output of the first computing unit, that is, the first frequency domain feature of the i-th audio frame, into the second computing unit. In addition, the first terminal also inputs at least one of the second frequency domain feature of the i-th audio frame and the second frequency domain feature of the (i + 2)-th audio frame into the second computing unit. The second computing unit fuses the first frequency domain feature of the i-th audio frame and at least one of the second frequency domain feature of the i-th audio frame and the second frequency domain feature of the (i + 2)-th audio frame, multiplies the fused information by the second weight matrix to obtain the first frequency domain feature of the (i + 1)-th audio frame.

[0123] That is to say, when extracting features from multiple second frequency domain features through the second feature extraction layer, instead of extracting features from the second frequency domain feature of the target audio frame, the second frequency domain feature of the target audio frame and the second frequency domain feature of the reference audio frame are combined for feature extraction, and the features obtained in this way can fully express the context information.

[0124] In addition, for the above-mentioned voice enhancement model, the second feature extraction layer can fully utilize the information of the first feature extraction layer. The more the number of feature extraction layers, the more context information of the frequency domain features obtained, and the higher the accuracy of the model. Of course, increasing the number of feature extraction layers will also increase the overhead, and the number of feature extraction layers is set by technicians according to the actual situation. For different usage scenarios and different inputs, appropriately increasing or decreasing the number of computing units or network levels can further reduce the computational complexity or increase the algorithm accuracy respectively. The relevant structures are as Figure 5 shown. Figure 5 It is an algorithm framework diagram with an increased hierarchical structure. In the figure, the computing unit can not only use GRU, but also use LSTM, or directly use the native RNN network. It is also possible to use a common RNN network in the first layer, GRU in the second layer, LSTM in the third layer, and others in the fourth layer. Each layer can use different network types in different orders. The input at a certain moment of the current layer can use the output at the corresponding moment of the previous layer and its outputs at the previous and next moments, or use the output at the corresponding moment of the previous layer and its outputs at the previous N moments and the next M moments. The more output information is utilized, the greater the computational amount, and the larger N and M are, the greater the network latency. Both N and M are positive integers.

[0125] See Figure 4 , the first terminal obtains the first frequency domain feature of the audio frame through the second feature extraction layer 403 based on the second frequency domain feature of the audio frame and the second frequency domain features of at least one audio frame adjacent to the audio frame. Figure 4 W2 in

[0126] is also the second weight matrix. In some embodiments, the above-mentioned computing unit is GRU. When the input speech uses the same sampling rate, the voice enhancement model can significantly reduce the computational amount while obtaining a denoising effect similar to that of other models.

[0127] 304. The first terminal obtains the first mask of the voice data based on the multiple first frequency domain features, and the first mask is used to remove the noise in the voice data.

[0128] In a possible implementation manner, the first terminal performs a fully connected process on the multiple first frequency domain features through the voice enhancement model to obtain the first mask of the voice data. In some embodiments, the first mask is also referred to as the prediction mask (Mask Pred).

[0129] For example, the first terminal multiplies the multiple first frequency domain features by the third weight matrix to obtain the first mask of the voice data. The third weight matrix is the weight matrix corresponding to the fully connected layer (output layer) of the voice enhancement model. See Figure 4, the first terminal inputs multiple first frequency-domain features into the output layer 404 of the speech enhancement model, and performs a fully-connected process on the multiple first frequency-domain features through the third weight matrix corresponding to the output layer 404 to obtain the first mask matrix of the speech data.

[0130] For example, the first terminal generates a frequency-domain feature matrix based on multiple frequency-domain features, and each row of the frequency-domain feature matrix is a frequency-domain feature. The first terminal performs a fully-connected process on the frequency-domain feature matrix through the speech enhancement model, that is, multiplies the frequency-domain feature matrix by the third weight matrix to obtain the first mask matrix of the speech data. Taking the frequency-domain feature matrix as The third weight matrix is as an example, the first terminal passes the frequency-domain feature matrix and the third weight matrix to multiply and obtain the first mask matrix of the speech data

[0131] 305. The first terminal generates target speech data based on the spectrum of the speech data and the first mask.

[0132] In a possible implementation manner, the first terminal multiplies multiple frequency points of the spectrum of the speech data by the first mask through the speech enhancement model to obtain a first target spectrum. The first terminal converts the first target spectrum into the target speech data.

[0133] Wherein, the spectrum of the speech data is composed of the frequency-domain information of multiple audio frames, and the multiple frequency points correspond one-to-one with the multiple initial frequency-domain information.

[0134] For example, the first terminal multiplies multiple frequency points of the spectrum of the speech data by the first mask matrix through the speech enhancement model to obtain a first target spectrum. The first terminal performs an inverse time-frequency transform on the first target spectrum to obtain the target speech data. For example, taking the spectrum of the speech data as The first mask matrix of the speech data is as an example, the first terminal passes the spectrum of the speech data of multiple frequency points and the first mask matrix to multiply and obtain the first target spectrum The first terminal converts the first target spectrum into the target speech data through the inverse Fourier transform. For example, the first terminal uses the Inverse Short-Time Fourier Transform (ISTFT) to convert the first target spectrum into the target speech data.

[0135] Of course, in addition to using the inverse short-time Fourier transform to convert the first target frequency spectrum into target speech data, the first terminal can also convert the first target frequency spectrum into target speech data by other means, which is not limited in the embodiments of the present application.

[0136] It should be noted that in the above steps 301-305, the first terminal is taken as an example of the execution entity for illustration. In other possible embodiments, the above steps 301-305 can also be executed by the server. In this case, when the server obtains the speech data through step 301, the speech data can be sent to the server by the first terminal or obtained by the server from the audio file, which is not limited in the embodiments of the present application.

[0137] The following will be combined with Figure 6 to illustrate the above steps 301-305. Refer to Figure 6 , the first terminal obtains speech data, the first terminal frames and windows the speech data to obtain a plurality of audio frames. The first terminal performs time-frequency transformation on the plurality of audio frames by using the short-time Fourier transform (STFT) to obtain a plurality of initial frequency domain information respectively corresponding to the plurality of audio frames. The first terminal performs normalization (Norm) processing on the plurality of initial frequency domain information to obtain a plurality of frequency domain information. The first terminal inputs the plurality of frequency domain information into a speech enhancement model (RNN) and obtains a first mask (MaskPred) of the speech data through the speech enhancement model. The first terminal multiplies the plurality of initial frequency domain information with the first mask correspondingly to obtain a first target frequency spectrum. The first terminal processes the first target frequency spectrum by using the inverse transform (ISTFT) of the short-time Fourier transform to obtain target speech data.

[0138] Optionally, after step 305, the first terminal can further execute the following step 306.

[0139] 306. The first terminal sends the target speech data to the server, and the server sends the target speech data to the second terminal.

[0140] It should be noted that if the above steps 301-305 are executed by the server, then after step 305, the server can directly send the target speech data to the second terminal.

[0141] The above steps 301-306 are described by taking the first terminal sending speech data to the second terminal as an example. When the second terminal sends speech data to the first terminal, the steps executed by the second terminal and the steps executed by the first terminal in the above steps 301-306 belong to the same inventive concept and will not be elaborated here.

[0142] Through the technical solution provided by the embodiments of the present application, when performing speech noise reduction, there is no need to identify noise through a model with a complex structure. Instead, a first mask is directly determined based on the frequency-domain information of the speech data, and the first mask is combined with the spectrum of the speech data to obtain the target speech data. While ensuring the noise reduction effect, the speed of speech noise reduction is improved, and the consumption of computing resources is reduced.

[0143] In the above description process, a speech enhancement model is involved. To understand the present application more clearly, the training method of the speech enhancement model will be described below. See Figure 7 The method includes:

[0144] 701. The terminal obtains first sample speech data and second sample speech data, and the second sample speech data is the speech data obtained by adding noise to the first sample speech data.

[0145] In some embodiments, the first sample speech data is also referred to as clean speech, and the second sample speech data is also referred to as noisy speech.

[0146] In a possible implementation manner, the terminal sends a sample data acquisition request to the server. The sample data acquisition request carries an identifier of the sample data, and the identifier of the sample data is used to indicate the type of the sample data. After receiving the sample data acquisition request, the server obtains the identifier of the sample data from the sample data acquisition request. The server queries in the sample database based on the identifier of the sample data to obtain a sample speech data set, which includes multiple first sample speech data and multiple second sample speech data, and the multiple first sample speech data and the multiple second sample speech data correspond one by one. The multiple first sample speech data and the multiple second sample speech data corresponding one by one means that each first sample speech data corresponds to one second sample speech data.

[0147] In a possible implementation manner, the terminal obtains multiple first sample speech data, and the first sample speech data is the speech data recorded in a noise-free environment. The terminal performs a noise addition process on the multiple first sample speech data to obtain multiple second sample speech data.

[0148] 702. The terminal inputs the second sample speech data into the speech enhancement model, and through the speech enhancement model, obtains a predicted first mask of the second sample speech data.

[0149] Among them, the method by which the terminal obtains the predicted first mask of the second sample speech data through the speech enhancement model belongs to the same inventive concept as the above steps 302-304. The implementation structure refers to the relevant descriptions of the above steps 302-305 and will not be elaborated here.

[0150] 703. The terminal multiplies multiple frequency points of the spectrum of the second sample voice data by the predicted first mask to obtain a predicted spectrum.

[0151] 704. The terminal trains the voice enhancement model based on the difference information between the predicted spectrum and the spectrum of the first sample voice data.

[0152] In a possible implementation manner, the terminal constructs a loss function based on the difference information between the predicted spectrum and the spectrum of the first sample voice data. The terminal trains the voice enhancement model based on the loss function.

[0153] For example, the terminal constructs a Magnitude Signal Approximation (MSA) loss function based on the difference information between the predicted spectrum and the spectrum of the first sample voice data, that is, directly makes a difference feedback network between the spectrum of the clean voice data and the spectrum of the processed noisy voice data, and then corrects the predicted first mask. The MSA loss function is the following formula (3).

[0154]

[0155] Where, L is the loss function, l and k respectively represent time and frequency, S[] is the clean voice data, that is, the label, X[] is the noisy voice data, that is, the input of the trained voice enhancement model, Pred is the first mask, that is, Mask Pred, which is a decimal between 0 and 1, and a is a constant.

[0156] It should be noted that in the above steps 701-704, the terminal is taken as the execution subject for illustration. In other possible implementation manners, the above steps 701-704 can also be executed by the server, and the embodiments of the present application do not make any limitations in this regard.

[0157] In the above steps 301-306, the voice data is denoised by the voice enhancement model to obtain the target voice. This method can also be referred to as a neural network-based denoising method. In addition to the above steps 301-306, the embodiments of the present application also provide another voice processing method. In this voice processing method, two denoising methods of neural network denoising and statistical denoising are combined. Taking the first terminal sending voice data to the second terminal as an example, see Figure 8 , the method includes:

[0158] 801. The first terminal obtains voice data.

[0159] Among them, step 801 and the above step 301 belong to the same inventive concept. For the implementation process, see the description of the above step 301, and details will not be repeated here.

[0160] 802. The first terminal obtains multiple frequency-domain information of the voice data, and the multiple frequency-domain information corresponds one-to-one to multiple audio frames of the voice data.

[0161] Among them, step 802 and the above-mentioned step 302 belong to the same inventive concept. For the implementation process, refer to the description of the above-mentioned step 302 and will not be elaborated here.

[0162] 803. The first terminal extracts features from the multiple frequency-domain information to obtain multiple first frequency-domain features, and the first frequency-domain features are determined based on the corresponding audio frames and the audio frames adjacent to the audio frames.

[0163] Among them, step 803 and the above-mentioned step 303 belong to the same inventive concept. For the implementation process, refer to the description of the above-mentioned step 303 and will not be elaborated here.

[0164] 804. The first terminal obtains a first mask of the voice data based on the multiple first frequency-domain features, and the first mask is used to remove noise from the voice data.

[0165] Among them, step 804 and the above-mentioned step 304 belong to the same inventive concept. For the implementation process, refer to the description of the above-mentioned step 304 and will not be elaborated here.

[0166] 805. The first terminal performs statistical noise reduction on the voice data to obtain a second mask of the voice data.

[0167] Among them, statistical noise reduction refers to a method of removing noise from voice data by utilizing the statistical stationarity of the noise.

[0168] In a possible implementation manner, the first terminal obtains a noise estimation spectrum of the voice data. The first terminal obtains the second mask based on the spectrum of the voice data and the noise estimation spectrum. In some embodiments, this implementation manner is also referred to as the spectral subtraction method.

[0169] For example, the first terminal performs voice activity detection (VAD) on the voice data to obtain noise data from the voice data. The first terminal performs time-frequency transformation on the noise data to obtain a noise estimation spectrum corresponding to the noise data. The first terminal subtracts the spectrum of the noise data from the spectrum of the voice data to obtain a voice estimation spectrum. The first terminal divides each frequency point in the voice estimation spectrum by the corresponding frequency point in the spectrum of the voice data to obtain a second mask of the voice data. Compared with the voice, the noise is relatively stable. In a voice data, the variance of the voice corresponding segment is large, and the variance of the noise corresponding segment is small. The first terminal can perform voice activity detection on the voice data based on the change of the variance, and obtain noise data from the voice data. For example, the first terminal divides the voice data into multiple voice segments, and obtains the variances of the multiple voice segments respectively. In response to the variance of any voice segment being less than or equal to the variance threshold, the voice segment is determined as a noise segment, that is, noise data.

[0170] In a possible implementation manner, the first terminal performs Wiener filtering on the voice data to obtain the second mask.

[0171] For example, the first terminal inputs the voice data into a Wiener filter, and filters the voice data through the Wiener filter to obtain the second mask. Among them, the Wiener filter is trained based on clean voice data and noisy voice data, and has the ability to reduce noise for voice data. The noisy voice data is the data obtained by adding noise to the clean voice data. The noise reduction here refers to multiplying the determined second mask by the voice data. When training the Wiener filter, the loss function is the mean square error between the clean voice and the noisy voice, and the training objective is to minimize the expectation of the mean square error.

[0172] It should be noted that the first terminal can obtain the second mask through any of the above implementation manners, and the embodiments of the present application do not make limitations in this regard.

[0173] 806. The first terminal generates the target voice data based on the second mask, the spectrum of the target voice data, and the first mask.

[0174] In a possible implementation manner, the first terminal fuses the first mask and the second mask to obtain a fused mask. The first terminal multiplies the fused mask by multiple frequency points of the spectrum of the target voice data to obtain a second target spectrum. The first terminal converts the second spectrum into the target voice data.

[0175] The method for the first terminal to obtain the second target spectrum will be described below.

[0176] In a possible implementation, the first mask and the second mask are two matrices of the same size, and the values at the same positions in the two matrices correspond to the same frequency point. When the first terminal fuses the first mask and the second mask, it is to fuse the two matrices into one matrix. When fusing the two matrices, for the values at the same positions in the two matrices, the first terminal retains the smaller value and deletes the larger value.

[0177] For example, for the two matrices and when fusing the first value in the upper left corner of the two matrices, since 70 < 80, the first terminal retains the value 70, and the finally obtained matrix is This matrix is also the matrix corresponding to the fused mask. The first terminal multiplies the matrix corresponding to the fused mask by multiple frequency points of the spectrum of the target voice data to obtain the second target spectrum. The first terminal converts the second target spectrum into the target voice data through inverse Fourier transform. For example, the first terminal uses the Inverse Short-Time Fourier Transform (ISTFT) to convert the second target spectrum into the target voice data.

[0178] In a possible implementation, the first mask and the second mask are two matrices of the same size, and the values at the same positions in the two matrices correspond to the same frequency point. When the first terminal fuses the first mask and the second mask, it is to fuse the two matrices into one matrix. When fusing the two matrices, for the values at the same positions in the two matrices, the first terminal retains the larger value and deletes the smaller value.

[0179] For example, for the two matrices and when fusing the first value in the upper left corner of the two matrices, since 70 < 80, the first terminal retains the value 80, and the finally obtained matrix is This matrix is also the matrix corresponding to the fused mask. The first terminal multiplies the matrix corresponding to the fused mask by multiple frequency points of the spectrum of the target voice data to obtain the second target spectrum.

[0180] In a possible implementation, the first mask and the second mask are two matrices of the same size, and the values at the same positions in the two matrices correspond to the same frequency point. When the first terminal fuses the first mask and the second mask, that is, fuses the two matrices into one matrix. When fusing the two matrices, for the values at the same positions in the two matrices, the first terminal adds the two values and then multiplies the result by a preset weight to obtain a fused mask. The preset weight is set by the technician according to the actual situation, such as set to 0.3, 0.5 or 0.6, etc., and the embodiments of the present application do not limit this.

[0181] For example, for the two matrices and when fusing the first value in the upper left corner of the two matrices, the first terminal adds 70 and 80 and then multiplies the result by the preset weight. If the preset weight is 0.5, then the value 75 is obtained, and the final obtained matrix is This matrix is also the matrix corresponding to the fused mask. The first terminal multiplies the matrix corresponding to the fused mask by multiple frequency points of the spectrum of the target voice data to obtain a second target spectrum.

[0182] It should be noted that when obtaining the second target spectrum, the first terminal can adopt any of the above methods, and the embodiments of the present application do not limit this.

[0183] Next, a method for the first terminal to convert the second spectrum into the target voice data will be described.

[0184] In a possible implementation, the first terminal converts the second target spectrum into the target voice data through inverse Fourier transform. For example, the first terminal uses the Inverse Short-Time Fourier Transform (ISTFT) to convert the second target spectrum into the target voice data.

[0185] It should be noted that in the above steps 801-806, the first terminal is taken as the execution subject for illustration. In other possible implementations, the above steps 801-806 can also be executed by the server. In this case, when the server obtains the voice data through step 801, the voice data can be sent to the server by the first terminal or obtained by the server from the audio file, and the embodiments of the present application do not limit this.

[0186] In the above steps 801-806, two methods of statistical noise reduction and neural network noise reduction are integrated. During use, the user can independently choose to use statistical noise reduction alone, neural network noise reduction alone, or a method that combines statistical noise reduction and neural network noise reduction simultaneously for speech enhancement. The embodiments of the present application do not limit this. The following will be combined with Figure 9 , to illustrate the above steps 801-806. Refer to Figure 9 . The first terminal obtains speech data, and the first terminal frames and windows the speech data to obtain multiple audio frames. The first terminal performs time-frequency transformation on the multiple audio frames using the short-time Fourier transform (STFT) to obtain multiple initial frequency domain information respectively corresponding to the multiple audio frames. The first terminal performs normalization (Norm) processing on the multiple initial frequency domain information to obtain multiple frequency domain information. The first terminal inputs the multiple frequency domain information into a speech enhancement model (RNN) and obtains a first mask (Mask Pred) of the speech data through the speech enhancement model. After the first terminal obtains multiple initial frequency domain information respectively corresponding to the multiple audio frames, it performs statistical noise reduction based on the multiple initial frequency domain information to obtain a second mask. The first terminal fuses the first mask and the second mask to obtain a fusion mask (Gain Fusion). The first terminal multiplies the fusion mask with the multiple initial frequency domain information correspondingly to obtain a second target spectrum. The first terminal processes the second target spectrum using the inverse transform (ISTFT) of the short-time Fourier transform to obtain target speech data.

[0187] Optionally, if the speech processing method provided by the embodiments of the present application is applied in the scenario of a voice call, then the first terminal further performs the following step 807.

[0188] 807. The first terminal sends the target speech data to the server, and the server sends the target speech data to the second terminal.

[0189] The above steps 801-807 are described by taking the first terminal sending speech data to the second terminal as an example. When the second terminal sends speech data to the first terminal, the steps performed by the second terminal are within the same inventive concept as the steps performed by the first terminal in the above steps 801-807, and will not be elaborated here.

[0190] Through the technical solution provided by the embodiments of the present application, when performing speech noise reduction, there is no need to identify noise through a model with a complex structure. Instead, a first mask and a second mask are directly determined based on the frequency domain information of the speech data. The first mask and the second mask are fused to obtain a fusion mask. This process is also the fusion of statistical noise reduction and neural network noise reduction. Combining the first mask with the spectrum of the speech data improves the speed of speech noise reduction while ensuring the noise reduction effect and reducing the consumption of computing resources.

[0191] Figure 10 is a schematic structural diagram of a voice processing device provided by an embodiment of the present application. Refer to Figure 10 , the device includes: a frequency-domain information acquisition module 1001, a feature extraction module 1002, a first mask acquisition module 1003, and a voice data generation module 1004.

[0192] The frequency-domain information acquisition module 1001 is configured to acquire multiple frequency-domain information of the voice data, and the multiple frequency-domain information corresponds one-to-one to multiple audio frames of the voice data.

[0193] The feature extraction module 1002 is configured to perform feature extraction on the multiple frequency-domain information to obtain multiple first frequency-domain features, and the first frequency-domain features are determined based on the corresponding audio frames and the audio frames adjacent to the audio frames.

[0194] The first mask acquisition module 1003 is configured to acquire a first mask of the voice data based on the multiple first frequency-domain features, and the first mask is used to remove noise in the voice data.

[0195] The voice data generation module 1004 is configured to generate target voice data based on the spectrum of the voice data and the first mask.

[0196] In a possible implementation manner, the frequency-domain information acquisition module 1001 is configured to perform frame division and windowing on the voice data to obtain the multiple audio frames. Perform time-frequency transformation on the multiple audio frames to obtain multiple initial frequency-domain information. Perform normalization processing on the multiple initial frequency-domain information to obtain the multiple frequency-domain information.

[0197] In a possible implementation manner, the frequency-domain information acquisition module 1001 is configured to acquire the mean and variance of the multiple initial frequency-domain information. Use the mean and the variance to perform normalization processing on the multiple initial frequency-domain information to obtain the multiple frequency-domain information.

[0198] In a possible implementation manner, the feature extraction module 1002 is configured to input the multiple frequency-domain information into a voice enhancement model, and perform feature extraction on the multiple frequency-domain information through the voice enhancement model to obtain multiple second frequency-domain features. Through the voice enhancement model, in accordance with the arrangement order of the multiple audio frames, based on the multiple second frequency-domain features, obtain the multiple first frequency-domain features.

[0199] In a possible implementation manner, for any one of the multiple audio frames, the feature extraction module 1002 is configured to obtain the first frequency-domain feature of the audio frame based on the second frequency-domain feature of the audio frame and the second frequency-domain features of at least one audio frame adjacent to the audio frame.

[0200] In a possible implementation, the first mask acquisition module 1003 is configured to perform a fully connected process on the multiple first frequency domain features through a voice enhancement model to obtain the first mask of the voice data.

[0201] In a possible implementation, the voice data generation module 1004 is configured to multiply multiple frequency points of the spectrum of the voice data by the first mask through a voice enhancement model to obtain a first target spectrum. Convert the first target spectrum into the target voice data.

[0202] In a possible implementation, the apparatus further includes:

[0203] A model training module, configured to obtain first sample voice data and second sample voice data, where the second sample voice data is the voice data obtained by adding noise to the first sample voice data. Input the second sample voice data into the voice enhancement model, and through the voice enhancement model, obtain a predicted first mask of the second sample voice data. Multiply multiple frequency points of the spectrum of the second sample voice data by the predicted first mask to obtain a predicted spectrum. Train the voice enhancement model based on the difference information between the predicted spectrum and the spectrum of the first sample voice data.

[0204] In a possible implementation, the apparatus further includes:

[0205] A second mask acquisition module, configured to perform statistical noise reduction on the voice data to obtain the second mask of the voice data.

[0206] The voice data generation module 1004 is further configured to generate the target voice data based on the second mask, the spectrum of the standard voice data, and the first mask.

[0207] In a possible implementation, the second mask acquisition module is configured to perform any one of the following:

[0208] Obtain the noise estimation spectrum of the voice data. Based on the noise estimation spectrum and the spectrum of the voice data, obtain the second mask.

[0209] Perform Wiener filtering on the voice data to obtain the second mask.

[0210] In a possible implementation, the voice data generation module 1004 is further configured to fuse the first mask and the second mask to obtain a fused mask. Multiply the fused mask by multiple frequency points of the spectrum of the target voice data to obtain a second target spectrum. Convert the second spectrum into the target voice data.

[0211] It should be noted that: when enhancing the voice by the voice enhancement device provided in the above embodiments, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the voice processing device provided in the above embodiments and the embodiments of the voice processing method belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.

[0212] Through the technical solution provided by the embodiments of the present application, when performing voice noise reduction, there is no need to identify noise through a model with a complex structure. Instead, a first mask is directly determined based on the frequency domain information of the voice data, and the first mask is combined with the spectrum of the voice data to obtain the target voice data. While ensuring the noise reduction effect, the speed of voice noise reduction is improved, and the consumption of computing resources is reduced.

[0213] The embodiments of the present application provide a computer device for executing the above method. The computer device can be implemented as a terminal or a server. Here, the terminal is also the above-mentioned first terminal or second terminal. First, the structure of the terminal will be introduced below:

[0214] Figure 11 It is a schematic structural diagram of a terminal provided by the embodiments of the present application. The terminal 1100 can be: a smart phone, a tablet computer, a notebook computer or a desktop computer. The terminal 1100 may also be referred to by other names such as a terminal, a portable terminal, a laptop terminal, a desktop terminal, etc.

[0215] Generally, the terminal 1100 includes: one or more processors 1101 and one or more memories 1102.

[0216] The processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1101 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1101 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0217] The memory 1102 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1102 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1102 is used to store at least one computer program, and the at least one computer program is used to be executed by the processor 1101 to implement the voice processing method provided in the method embodiments of the present application.

[0218] In some embodiments, the terminal 1100 may further optionally include: a peripheral device interface 1103 and at least one peripheral device. The processor 1101, the memory 1102, and the peripheral device interface 1103 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1103 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, a positioning assembly 1108, and a power supply 1109.

[0219] The peripheral device interface 1103 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 can be implemented on separate chips or circuit boards, and this embodiment does not limit this.

[0220] The radio frequency circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1104 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1104 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on.

[0221] The display screen 1105 is used to display a UI (User Interface). The UI can include graphics, text, icons, videos, and any combination thereof. When the display screen 1105 is a touch display screen, the display screen 1105 also has the ability to collect touch signals on or above the surface of the display screen 1105. The touch signal can be input to the processor 1101 as a control signal for processing. At this time, the display screen 1105 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard.

[0222] The camera component 1106 is used to collect images or videos. Optionally, the camera component 1106 includes a front camera and a rear camera. Generally, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal.

[0223] The audio circuit 1107 can include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1101 for processing, or input to the radio frequency circuit 1104 to achieve voice communication.

[0224] The positioning component 1108 is used to locate the current geographical location of the terminal 1100 to implement navigation or LBS (Location-Based Service).

[0225] The power supply 1109 is used to supply power to each component in the terminal 1100. The power supply 1109 can be alternating current, direct current, a primary battery, or a rechargeable battery.

[0226] In some embodiments, the terminal 1100 further includes one or more sensors 1110. The one or more sensors 1110 include but are not limited to: an acceleration sensor 1111, a gyroscope sensor 1112, a pressure sensor 1113, a fingerprint sensor 1114, an optical sensor 1115, and a proximity sensor 1116.

[0227] The acceleration sensor 1111 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal 1100.

[0228] The gyroscope sensor 1112 can detect the body orientation and rotation angle of the terminal 1100. The gyroscope sensor 1112 can cooperate with the acceleration sensor 1111 to collect the 3D actions of the user on the terminal 1100.

[0229] The pressure sensor 1113 can be disposed on the side frame of the terminal 1100 and / or under the display screen 1105. When the pressure sensor 1113 is disposed on the side frame of the terminal 1100, it can detect the holding signal of the user on the terminal 1100, and the processor 1101 can perform left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor 1113. When the pressure sensor 1113 is disposed under the display screen 1105, the processor 1101 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1105.

[0230] The fingerprint sensor 1114 is used to collect the fingerprint of the user. The processor 1101 can identify the user's identity according to the fingerprint collected by the fingerprint sensor 1114, or the fingerprint sensor 1114 can identify the user's identity according to the collected fingerprint.

[0231] The optical sensor 1115 is used to collect the ambient light intensity. In one embodiment, the processor 1101 can control the display brightness of the display screen 1105 according to the ambient light intensity collected by the optical sensor 1115.

[0232] The proximity sensor 1116 is used to collect the distance between the user and the front of the terminal 1100.

[0233] Those skilled in the art can understand that Figure 11 the structure shown in [[ ]] does not constitute a limitation on the terminal 1100, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.

[0234] The above computer device can also be implemented as a server. The structure of the server will be introduced below:

[0235] Figure 12 FIG. is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1200 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 1201 and one or more memories 1202. Among them, at least one computer program is stored in the one or more memories 1202, and the at least one computer program is loaded and executed by the one or more processors 1201 to implement the methods provided in the above various method embodiments. Of course, the server 1200 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server 1200 may also include other components for implementing the functions of the device, which will not be elaborated here.

[0236] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program. The above computer program can be executed by a processor to complete the voice processing method in the above embodiment. For example, the computer-readable storage medium may be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0237] In an exemplary embodiment, a computer program product or a computer program is also provided. The computer program product or the computer program includes program code, and the program code is stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the above voice processing method.

[0238] In some embodiments, the computer program involved in the embodiments of the present application may be deployed to be executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected through a communication network. The multiple computer devices distributed at multiple locations and interconnected through a communication network may form a blockchain system.

[0239] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a disk, an optical disc, etc.

[0240] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A voice processing method, characterized in that, The method includes: Obtaining multiple frequency domain information of the speech data, where the multiple frequency domain information corresponds one-to-one to multiple audio frames of the speech data; Inputting the multiple frequency domain information into a speech enhancement model, and through a first feature extraction layer of the speech enhancement model, sequentially performing feature extraction on the multiple frequency domain information, combining the result of the previous feature extraction each time during feature extraction to obtain multiple second frequency domain features; through a second feature extraction layer of the speech enhancement model, according to the arrangement order of the multiple audio frames, based on the multiple second frequency domain features, obtaining multiple first frequency domain features, where the first frequency domain features are determined based on the corresponding audio frame and the audio frames adjacent to the audio frame; Based on the multiple first frequency domain features, obtaining a first mask of the speech data, where the first mask is used to remove noise in the speech data; Generating target speech data based on the spectrum of the speech data and the first mask.

2. The method according to claim 1, wherein The obtaining of the multiple frequency domain information of the speech data includes: Framing and windowing the speech data to obtain the multiple audio frames; Performing time-frequency transformation on the multiple audio frames to obtain multiple initial frequency domain information; Performing normalization processing on the multiple initial frequency domain information to obtain the multiple frequency domain information.

3. The method according to claim 2, characterized in that, The performing of the normalization processing on the multiple initial frequency domain information to obtain the multiple frequency domain information includes: Obtaining the mean and variance of the multiple initial frequency domain information; Using the mean and the variance to perform normalization processing on the multiple initial frequency domain information to obtain the multiple frequency domain information.

4. The method according to claim 1, characterized in that The obtaining of the multiple first frequency domain features according to the arrangement order of the multiple audio frames and based on the multiple second frequency domain features includes: For any one of the multiple audio frames, based on the second frequency domain feature of the audio frame and the second frequency domain features of at least one audio frame adjacent to the audio frame, obtaining the first frequency domain feature of the audio frame.

5. The method according to claim 1, characterized in that, The obtaining of the first mask of the speech data based on the multiple first frequency domain features includes: Through the speech enhancement model, performing a fully connected process on the multiple first frequency domain features to obtain the first mask of the speech data.

6. The method according to claim 1, wherein The generating of the target speech data based on the spectrum of the speech data and the first mask includes: Through the speech enhancement model, multiplying multiple frequency points of the spectrum of the speech data by the first mask to obtain a first target spectrum; Converting the first target spectrum into the target speech data.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Obtaining first sample speech data and second sample speech data, where the second sample speech data is the speech data obtained by adding noise to the first sample speech data; Inputting the second sample speech data into the speech enhancement model, and through the speech enhancement model, obtaining a predicted first mask of the second sample speech data; Multiplying multiple frequency points of the spectrum of the second sample speech data by the predicted first mask to obtain a predicted spectrum; Training the speech enhancement model based on the difference information between the predicted spectrum and the spectrum of the first sample speech data.

8. The method according to claim 1, wherein Before generating the target speech data based on the spectrum of the speech data and the first mask, the method further includes: Performing statistical noise reduction on the speech data to obtain a second mask of the speech data; The generating the target speech data based on the spectrum of the speech data and the first mask includes: Generating the target speech data based on the second mask, the spectrum of the speech data, and the first mask.

9. The method according to claim 8, characterized in that, The performing statistical noise reduction on the speech data to obtain a second mask of the speech data includes any one of the following: Obtaining a noise estimation spectrum of the speech data; obtaining the second mask based on the noise estimation spectrum and the spectrum of the speech data; Performing Wiener filtering on the speech data to obtain the second mask.

10. The method according to claim 8, wherein The generating the target speech data based on the second mask, the spectrum of the speech data, and the first mask includes: Fusing the first mask and the second mask to obtain a fused mask; Multiplying the fused mask by multiple frequency points of the spectrum of the speech data to obtain a second target spectrum; Converting the second target spectrum into the target speech data.

11. A voice processing device, characterized in that, The apparatus includes: A frequency domain information acquisition module, configured to acquire multiple frequency domain information of speech data, where the multiple frequency domain information corresponds one-to-one to multiple audio frames of the speech data; A feature extraction module, configured to input the multiple frequency domain information into a speech enhancement model, and sequentially perform feature extraction on the multiple frequency domain information through a first feature extraction layer of the speech enhancement model. Each time feature extraction is performed, the result of the previous feature extraction is combined to obtain multiple second frequency domain features; through a second feature extraction layer of the speech enhancement model, based on the multiple second frequency domain features, multiple first frequency domain features are obtained according to the arrangement order of the multiple audio frames, and the first frequency domain features are determined based on the corresponding audio frame and the audio frames adjacent to the audio frame; A first mask acquisition module, configured to acquire a first mask of the speech data based on the multiple first frequency domain features, where the first mask is used to remove noise in the speech data; A speech data generation module, configured to generate target speech data based on the spectrum of the speech data and the first mask.

12. The device according to claim 11, wherein The frequency domain information acquisition module is configured to frame and window the speech data to obtain the multiple audio frames; perform time-frequency transformation on the multiple audio frames to obtain multiple initial frequency domain information; Perform normalization processing on the multiple initial frequency domain information to obtain the multiple frequency domain information.

13. The device according to claim 12, characterized in that, The frequency domain information acquisition module is configured to obtain the mean and variance of the multiple initial frequency domain information; use the mean and the variance to perform normalization processing on the multiple initial frequency domain information to obtain the multiple frequency domain information.

14. The device according to claim 11, characterized in that, The feature extraction module is configured to, for any one of the multiple audio frames, obtain the first frequency domain feature of the audio frame based on the second frequency domain feature of the audio frame and the second frequency domain features of at least one audio frame adjacent to the audio frame.

15. The device according to claim 11, characterized in that, The first mask acquisition module is configured to perform a fully-connected process on the multiple first frequency-domain features through a voice enhancement model to obtain a first mask of the voice data.

16. The device according to claim 11, characterized in that, The voice data generation module is configured to multiply multiple frequency points of the spectrum of the voice data by the first mask through a voice enhancement model to obtain a first target spectrum; and convert the first target spectrum into the target voice data.

17. The device according to any one of claims 11-16, characterized in that, The device further includes: A model training module, configured to obtain first sample voice data and second sample voice data, where the second sample voice data is voice data obtained by adding noise to the first sample voice data; input the second sample voice data into the voice enhancement model, and through the voice enhancement model, obtain a predicted first mask of the second sample voice data; multiply multiple frequency points of the spectrum of the second sample voice data by the predicted first mask to obtain a predicted spectrum; and train the voice enhancement model based on the difference information between the predicted spectrum and the spectrum of the first sample voice data.

18. The device according to claim 11, characterized in that, The device further includes: A second mask acquisition module, configured to perform statistical noise reduction on the voice data to obtain a second mask of the voice data; The voice data generation module is further configured to generate the target voice data based on the second mask, the spectrum of the voice data, and the first mask.

19. The device according to claim 18, characterized in that, The second mask acquisition module is configured to perform any one of the following: Obtain a noise estimation spectrum of the voice data; and obtain the second mask based on the noise estimation spectrum and the spectrum of the voice data; Perform Wiener filtering on the voice data to obtain the second mask.

20. The device according to claim 18, characterized in that, The voice data generation module is further configured to fuse the first mask and the second mask to obtain a fused mask; multiply multiple frequency points of the spectrum of the voice data by the fused mask to obtain a second target spectrum; and convert the second target spectrum into the target voice data.

21. A computer device, characterized in that, The computer device includes one or more processors and one or more memories, and at least one computer program is stored in the one or more memories. The computer program is loaded and executed by the one or more processors to implement the voice processing method according to any one of claims 1 to 10.

22. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium. The computer program is loaded and executed by a processor to implement the voice processing method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Universal single channel real-time noise-reduction method

    CN107452389A

  • Voice call processing method and related device

    CN111179957A

  • Noise estimation method and system for far-field call

    CN111696567A

  • Audio noise reduction method and training method of audio noise reduction model

    CN111883091A