Voice signal processing method and device, electronic equipment and storage medium
Through the serial fusion method of CNN and GRU, the problem of high hardware performance and computing complexity of deep learning voice signal processing technology is solved, and efficient voice signal noise reduction and quality improvement under low complexity is achieved.
Patent Information
- Application Number
- CN202410009829.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
The existing deep learning voice signal processing technology has high hardware performance requirements and high computing complexity, which cannot meet the real-time computing requirements, resulting in a reduced voice signal noise reduction effect.
Using the serial fusion method of CNN and GRU, the feature extraction, multiple feature transformations, jump connection feature mapping and feature fusion of the enhanced voice signal is reduced, and the computational complexity is improved.
While reducing the computational complexity, it effectively removes noise, meets the noise reduction requirements of real-time communication, and improves voice signal quality and processing speed.
Smart Images

Figure CN120260591A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technologies, and in particular, to a method, apparatus, electronic device, and storage medium for processing voice signals. Background Art
[0002] Voice signal processing belongs to a branch of digital signal processing. Due to the continuous development of voice signal processing technology, the application fields involved in voice signal processing technology are very extensive, including voice calls, telephone conferences, scene recordings, hearing aid devices, and voice recognition devices, etc., and it has become a preprocessing module for many voice coding and recognition systems. Different from traditional digital signal processing methods, deep learning voice signal processing technology utilizes the powerful non-linear mapping ability of deep neural network structures. Through the training of a large amount of data, a non-linear model is trained to process voice signals.
[0003] In the related technologies of deep learning voice signal processing, a convolutional neural network (CNN), a long short-term memory neural network (LSTM), or a gated recurrent unit module (GRU) is usually used alone to process voice signals, or various unit modules are used in combination. However, these deep learning voice signal processing technologies all have high requirements for the performance of hardware devices and have high computational complexity, and cannot meet the requirements of real-time operation, resulting in a significant decrease in the noise reduction effect on voice signals. Summary of the Invention
[0004] Embodiments of this application provide a method, apparatus, and computer-readable storage medium for processing voice signals, which can improve the voice signal processing effect and the voice signal processing rate.
[0005] The technical solution of the embodiments of this application is implemented as follows:
[0006] Embodiments of this application provide a method for processing voice signals. The method includes: extracting features from a voice signal to be enhanced to obtain voice features to be enhanced; performing multiple feature transformations on the voice features to be enhanced in a loop to obtain a voice feature vector; performing feature mapping on the voice feature vector by using a skip connection method to obtain a voice mapping feature; and performing feature fusion on the voice feature vector and the voice mapping feature to obtain the gain information of the voice signal to be enhanced.
[0007] An embodiment of the present application provides a voice signal processing device, including: a feature extraction module for extracting features from the voice signal to be enhanced to obtain voice features to be enhanced; a feature transformation module for performing multiple feature transformations on the voice features to be enhanced in a loop to obtain a voice feature vector; a feature mapping module for performing feature mapping on the voice feature vector in a jump connection manner to obtain a voice mapped feature; a feature fusion module for performing feature fusion on the voice feature vector and the voice mapped feature to obtain the gain information of the voice signal to be enhanced.
[0008] In some embodiments, the feature transformation module is further configured to: perform multiple first feature transformations on the voice features to be enhanced in a loop through a first feature transformation module with a specific number of channels to obtain a first voice feature vector; wherein, during the multiple first feature transformations, the specific number of channels of the first feature transformation module increases successively; perform at least one attention feature transformation on the first voice feature vector through an attention module based on an attention mechanism to obtain a second voice feature vector; perform multiple second feature transformations on the second voice feature vector in a loop through a second feature transformation module with a specific number of channels to obtain the voice feature vector; wherein, during the multiple second feature transformations, the specific number of channels of the second feature transformation module decreases successively.
[0009] In some embodiments, the first feature transformation module includes a first convolutional layer, a first normalization layer, a first activation layer, and a first attention mechanism layer; the feature transformation module is further configured to: during each first feature transformation, perform convolutional processing on the voice features to be enhanced through the first convolutional layer to obtain a first convolutional feature; perform normalization processing on the first convolutional feature through the first normalization layer to obtain a first normalized feature; perform activation processing on the first normalized feature through the first activation layer to obtain a first activated feature; perform feature allocation on the first activated feature through the first attention mechanism layer to obtain the first voice feature vector.
[0010] In some embodiments, the second feature transformation module includes a second convolutional layer, a second normalization layer, a second activation layer, and a second attention mechanism layer; the specific number of channels of the first feature transformation module when performing the last first feature transformation on the voice features to be enhanced is the same as the specific number of channels of the second feature transformation module when performing the first second feature transformation on the second voice feature vector, and, the dimension of the voice feature vector obtained after performing the last second feature transformation on the second voice feature vector is the same as the dimension of the voice features to be enhanced.
[0011] In some embodiments, the feature mapping module is further configured to: perform feature mapping on the speech feature vector in a skip connection manner through a multi-layer gated recurrent unit having a skip connection relationship to obtain a speech mapping feature; wherein, the multi-layer gated recurrent unit having a skip connection relationship has N layers; the input feature of the first-layer gated recurrent unit is the speech feature vector; the input feature of the Nth-layer gated recurrent unit is the output features of all the gated recurrent units connected before the Nth-layer gated recurrent unit and the speech feature vector; N is a positive integer greater than 1.
[0012] In some embodiments, for the Nth-layer gated recurrent unit, the feature mapping module is further configured to: obtain the output features of all the gated recurrent units connected before the Nth-layer gated recurrent unit and the speech feature vector; perform a vector summation operation on the output features of all the gated recurrent units connected before the Nth-layer gated recurrent unit and the speech feature vector to obtain the input feature of the Nth-layer gated recurrent unit; perform feature mapping on the input feature through the Nth-layer gated recurrent unit to obtain the output feature of the Nth-layer gated recurrent unit.
[0013] In some embodiments, the apparatus further includes a signal enhancement module, and the signal enhancement module is configured to: perform band decompression on the gain information to obtain decompressed gain information; perform a multiplication operation on the decompressed gain information and the power spectrum of the speech audio feature to obtain an enhanced audio feature; perform an inverse frequency domain transformation on the enhanced audio feature to obtain an enhanced speech signal of the speech feature to be enhanced.
[0014] In some embodiments, the decompressed gain information includes a plurality of gain values; the signal enhancement module is further configured to: obtain the power value of the speech audio feature at each frequency component, and there is a preset mapping relationship between the power value at each frequency component and one gain value in the decompressed gain information; based on the preset mapping relationship, perform a multiplication operation on each gain value and the power value to obtain the enhanced audio feature.
[0015] In some embodiments, the feature extraction module is further configured to: perform a frequency domain transformation on the speech feature to be enhanced to obtain a speech audio feature; perform audio feature extraction on the speech audio feature to obtain an audio feature vector; perform band compression on the audio feature vector to obtain the speech feature to be enhanced.
[0016] In some embodiments, the voice signal processing method is implemented by a voice signal processing model; the device further includes a model training module, and the model training module is used to obtain sample data; the sample data includes a first type of voice sample, a second type of voice sample corresponding to the first type of voice sample, and a true gain between the first type of voice sample and the second type of voice sample; the first type of voice sample is a clean voice sample without noise, and the second type of voice sample is a voice sample obtained by adding noise to the first type of voice sample; perform feature preprocessing on the first type of voice sample and the second type of voice sample to obtain sample voice features to be enhanced, and input the sample voice features to be enhanced into the voice signal processing model; through the feature transformation network of the voice signal processing model, perform multiple feature transformations on the sample voice features to be enhanced in a loop to obtain sample voice feature vectors; through the feature mapping network of the voice signal processing model, perform feature mapping on the sample voice feature vectors in a jump connection manner to obtain sample voice mapping features; through the feature fusion network of the voice signal processing model, perform feature fusion on the sample voice feature vectors and the sample voice mapping features to obtain a sample estimated gain corresponding to the sample voice features to be enhanced; input the sample estimated gain and the true gain into a loss model for loss calculation to obtain a loss result; update the model parameters in the voice signal processing model based on the loss result to obtain a trained voice signal processing model.
[0017] In some embodiments, the model training module is further used for: respectively performing audio feature extraction on the first type of voice sample and the second type of voice sample to correspondingly obtain a first type of sample audio feature vector and a second type of sample audio feature vector; performing feature splicing on the first type of sample audio feature vector and the second type of sample audio feature vector to obtain sample spliced voice features; performing band compression on the sample spliced voice features to obtain the sample voice features to be enhanced.
[0018] An embodiment of the present application provides an electronic device, including: a memory for storing computer-executable instructions; a processor for implementing the voice signal processing method provided by the embodiment of the present application when executing the computer-executable instructions stored in the memory.
[0019] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for implementing the voice signal processing method provided by the embodiment of the present application when being executed by a processor.
[0020] An embodiment of the present application provides a computer program product, which includes executable instructions stored in a computer-readable storage medium. When a processor of an electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the voice signal processing method provided by the embodiment of the present application is implemented.
[0021] The embodiment of the present application has the following beneficial effects:
[0022] First, feature extraction is performed on the voice signal to be enhanced, and the voice features to be enhanced obtained after feature extraction are cyclically subjected to multiple feature transformations. Then, a jump connection method is used to perform feature mapping on the voice feature vectors obtained by the feature transformation. The voice feature vectors and the voice mapping features obtained by the feature mapping are subjected to feature fusion to obtain the gain information of the voice signal to be enhanced. Finally, the enhanced voice signal of the voice features to be enhanced is obtained through the signal enhancement process. In this way, by cyclically performing multiple feature transformations, the network depth and the number of parameters are increased, so that the voice features to be enhanced can be learned more effectively. And, through the jump connection in the feature mapping process, the continuity and correlation between the voice feature vectors can be ensured. In this way, after fusing the feature transformation results obtained after multiple feature transformations and the feature mapping results obtained after feature mapping, accurate gain information can be obtained, and the calculation rate of the gain information is improved, so as to ensure the processing effect on the voice signal to be enhanced during the enhanced signal processing process, and improve the enhanced signal processing rate and the voice signal quality. Description of the Drawings
[0023] Figure 1 is a schematic structural diagram of the voice signal processing system architecture provided by the embodiment of the present application;
[0024] Figure 2 is a schematic structural diagram of the voice signal processing device provided by the embodiment of the present application;
[0025] Figure 3 is an optional flowchart of the voice signal processing method provided by the embodiment of the present application;
[0026] Figure 4 is another optional flowchart of the voice signal processing method provided by the embodiment of the present application;
[0027] Figure 5 is a schematic diagram of the implementation process of obtaining the first voice feature vector provided by the embodiment of the present application;
[0028] Figure 6 is a schematic diagram of the implementation process of obtaining the voice feature vector provided by the embodiment of the present application;
[0029] Figure 7AIt is a schematic diagram of the implementation process of the feature mapping of the three-layer gated recurrent unit provided by the embodiments of the present application;
[0030] Figure 7B It is a schematic diagram of the implementation process of the feature mapping of the four-layer gated recurrent unit provided by the embodiments of the present application;
[0031] Figure 7C It is a schematic diagram of the implementation process of the feature mapping of the five-layer gated recurrent unit provided by the embodiments of the present application;
[0032] Figure 8 It is a schematic diagram of the process flow of the speech signal processing model training method provided by the embodiments of the present application;
[0033] Figure 9 It is a schematic diagram of the structure of the neural network model training provided by the embodiments of the present application;
[0034] Figure 10 It is a schematic diagram of the original speech signal provided by the embodiments of the present application;
[0035] Figure 11 It is a schematic diagram of the enhanced speech signal provided by the embodiments of the present application. Detailed implementation manners
[0036] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0037] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0038] If similar descriptions such as "first / second" appear in the application documents, the following description will be added. In the following description, the terms "first\second\third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0039] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0040] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0041] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.
[0042] 1) Ultra-wideband: The sampling rate is greater than or equal to 32 KHz, and the bandwidth is greater than or equal to 16 KHz (the bandwidth is 1 / 2 of the sampling rate).
[0043] 2) GRU network: A neural network based on GRU modules, which is a type of recurrent neural network (RNN). It controls the flow of information through a gating mechanism to model long-term dependencies in sequences and can solve problems such as the inability to remember for a long time and gradients in backpropagation in RNNs.
[0044] 3) CNN network: A neural network based on CNN modules, which is a feedforward neural network with convolutional calculations and a deep structure, and has the ability of feature learning, that is, it can extract high-order features from input information.
[0045] 4) Attention module: Also known as the attention mechanism module, it is a technology that enables the model to focus on important information, learn and absorb it fully. Generally speaking, it focuses the attention on important points and ignores other unimportant factors.
[0046] 5) Speech enhancement: When the speech signal is interfered with or even submerged by various noises, the technology of extracting useful speech signals from the noise background and suppressing and reducing noise interference.
[0047] Speech signal processing is a complex speech noise reduction technology that needs to remove environmental background noise from noisy speech signals and extract useful and clear speech signals. Moreover, to meet the speech quality requirements of real-time communication, how to improve the noise reduction effect on noisy speech signals based on low computational complexity has become a key challenge in speech signal processing.
[0048] In related technologies, various neural network units are usually used alone, or various units and modules are used in combination. However, these methods all have high computational complexity, and when these methods are applied to real-time communication, the speech signal processing process also has high requirements for hardware performance. On real-time communication software, these methods are difficult to use and cannot meet the real-time operation requirements. If the network structure of these methods is adjusted to meet the real-time operation requirements, the noise reduction effect on noisy speech signals will be greatly reduced. At the same time, although the computational complexity of the LSTM and GRU modules is low, the processing effect is not ideal.
[0049] Based on at least one of the above problems existing in related technologies, the embodiments of this application implement ultra-wideband speech signal processing based on the serial fusion of CNN and GRU. Compared with traditional large CNN models or hyperparameters and speech enhancement algorithms based on complex CNN models, the method proposed in the embodiments of this application has low computational complexity, and at the same time uses the temporal characteristics of GRU to further effectively remove noise, so as to meet the noise reduction effect requirements on real-time communication software.
[0050] The following describes the exemplary applications of the speech signal processing device (i.e., electronic device) provided in the embodiments of this application. The device provided in the embodiments of this application can be implemented as various types of user terminals capable of data processing or speech signal processing, such as laptops, tablets, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), smartphones, smart speakers, smart watches, smart TVs, in-vehicle terminals, etc., or can also be implemented as a server. Below, the exemplary application will be described when the speech signal processing device is implemented as a server.
[0051] See Figure 1 , Figure 1 FIG. is a schematic structural diagram of the architecture of the speech signal processing system 100 provided in the embodiments of this application. To implement a speech signal processing application, the speech information enhancement application runs on the terminal 400. The terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.
[0052] The terminal 400 is used to send a voice signal processing request to the server 200. The server 200 constitutes the voice signal processing device of the embodiment of the present application. The server 200 is used to respond to the voice signal processing request, obtain the voice signal to be enhanced, and perform feature extraction on the voice signal to be enhanced to obtain the voice feature to be enhanced; perform multiple feature transformations on the voice feature to be enhanced in a loop to obtain a voice feature vector; perform feature mapping on the voice feature vector in a jump connection manner to obtain a voice mapped feature; perform feature fusion on the voice feature vector and the voice mapped feature to obtain the gain information of the voice signal to be enhanced. After obtaining the gain information of the voice signal to be enhanced, the server 200 returns the gain information of the voice signal to be enhanced to the terminal 400, so as to output the gain information at the terminal 400 or perform signal enhancement on the voice feature to be enhanced based on the gain information to obtain the enhanced voice signal of the voice feature to be enhanced.
[0053] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal 400 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, and are not limited in the embodiment of the present application.
[0054] See Figure 2 , Figure 2 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Figure 2 The electronic device shown may be a voice signal processing device. The voice signal processing device includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the voice signal processing device is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. The bus system 440 includes not only a data bus, but also a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 2 all kinds of buses are labeled as the bus system 440.
[0055] The processor 410 can be an integrated circuit chip with the ability to process signals, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0056] The user interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.
[0057] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 450 optionally includes one or more storage devices that are physically located away from the processor 410.
[0058] The memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM), and the volatile memory can be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0059] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.
[0060] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0061] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wi-Fi (Wireless Fidelity), Universal Serial Bus (USB), etc.; A presentation module 453 is used to enable the presentation of information (such as a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (such as a display screen, a speaker, etc.); An input processing module 454 is used to detect one or more user inputs or interactions from one of the one or more input devices 432 and translate the detected inputs or interactions.
[0062] In some embodiments, the device provided by the embodiments of the present application can be implemented in software. Figure 2 A voice signal processing device 455 stored in the memory 450 is shown, which can be software in the form of a program and a plug-in, etc., and includes the following software modules: a feature extraction module 4551, a feature transformation module 4552, a feature mapping module 4553, a feature fusion module 4554, and a signal enhancement module 4555. These modules are logical, so they can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.
[0063] In other embodiments, the device provided by the embodiments of the present application can be implemented in hardware. As an example, the device provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the voice signal processing method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor can adopt one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Programmable Logic Devices (PLDs), Complex Programmable Logic Devices (CPLDs), Field-Programmable Gate Arrays (FPGAs), or other electronic components.
[0064] In some embodiments, a terminal or a server may implement the voice signal processing method provided in the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions may be microprogram-level commands, machine instructions, or software instructions. The computer program may be a native program or a software module in an operating system; it may be a native application (APP), that is, a program that needs to be installed in the operating system to run, or it may be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded to the browser environment to run. In short, the above computer-executable instructions may be instructions in any form, and the above computer program may be an application, module, or plug-in in any form.
[0065] The voice signal processing method provided in each embodiment of the present application may be executed by an electronic device, where the electronic device may be a server or a terminal, that is, the voice signal processing method in each embodiment of the present application may be executed by the server, or may be executed by the terminal, or may also be executed through the interaction between the server and the terminal.
[0066] See Figure 3 , Figure 3 is an optional flowchart of the voice signal processing method provided in the embodiments of the present application, and will be described in conjunction with Figure 3 the steps shown. Taking the execution subject of the voice signal processing method as the server as an example, the method includes the following steps S101 to step S104:
[0067] Step S101, perform feature extraction on the voice signal to be enhanced to obtain the voice features to be enhanced.
[0068] In the embodiments of the present application, the speech signal to be enhanced is a speech signal that needs to be enhanced in the speech signal processing service. Each speech signal that needs to be denoised can be used as the speech signal to be enhanced, that is, the speech signal to be enhanced is a noisy speech signal. The speech signal can be a sound signal collected in real time through a sound collection device such as a microphone, or a sound signal saved offline through a storage device. Speech signals are generally divided into clean speech signals and noisy speech signals. A clean speech signal is a speech signal that does not contain any noise other than the target speech in the speech signal; a noisy speech signal is a speech signal that contains other types of noise in addition to the target speech, that is, the target speech is interfered by noise and a clear speech signal cannot be obtained. For example, a voice call when there is another third person speaking nearby during a voice call, a conference call voice with the voices of other non-speakers during a conference call, a scene recording in different noisy scenarios, and a speech signal obtained in scenarios such as speech recognition, etc., all belong to noisy speech signals, that is, the speech signals to be enhanced. In real life, speech signals generally carry noise. Before further processing the signals (such as speech recognition, speech coding, etc.), it is often necessary to denoise the signals.
[0069] For the speech signal to be enhanced, the purpose of speech signal processing is to extract as pure a target speech as possible from the speech signal to be enhanced, improve the signal-to-noise ratio, and improve the speech quality. Currently, there are many methods for speech signal processing, including traditional signal processing methods and deep learning methods. The speech signal processing method provided in the embodiments of the present application belongs to the deep learning method. When performing speech signal processing, the speech signal to be enhanced can be preprocessed before being input into the network model. The preprocessing can be frequency domain transformation or frequency band compression of the speech signal to be enhanced, that is, feature extraction is performed on the speech signal to be enhanced to obtain the enhanced speech features after feature extraction, and the enhanced speech features are used as the input features of the network model to implement effective denoising processing using the network model.
[0070] For example, after the speech signal to be enhanced is collected through a microphone, according to the actual situation, first, the speech signal to be enhanced can be subjected to a fast Fourier transform (FFT) operation of N (such as 256, 512, 1024) points to obtain the real and imaginary parts of the frequency domain features corresponding to the speech signal to be enhanced. Subsequently, only the first N / 2 + 1 (such as 129, 257, 513) frequency domain features are used as the speech audio features. Then, the neural network layer can be used to extract the frequency domain features of the speech audio features. Finally, a frequency band compression method can be used for frequency band compression to reduce the feature dimension of the input features and obtain the enhanced speech features, thereby reducing the operation complexity of the subsequent network model.
[0071] Step S102: Perform multiple feature transformations on the voice feature to be enhanced in a loop to obtain a voice feature vector.
[0072] In the embodiment of the present application, through a feature transformation network, the voice feature to be enhanced is subjected to feature transformation. The feature transformation network may include different feature transformation modules, and multiple feature transformations are performed through different feature transformation modules. In the feature transformation module, the loop iteration times of the network can be set according to the actual situation. For example, loop 3 times, and the loop iteration times of the network are not limited. When the loop iteration times threshold of each feature transformation module is reached, the loop stops, and the output feature of the feature transformation module is input into the next feature transformation module until the feature transformation process is completed.
[0073] In some embodiments, the feature transformation module may be composed of a convolutional neural network. In each loop iteration process of performing feature transformation, the number of output channels of the network can be changed according to actual needs. By increasing or decreasing the number of output channels of the network, the network depth and the number of parameters of the feature transformation network are increased, and the input feature data can be learned more effectively.
[0074] Step S103: Perform feature mapping on the voice feature vector in a jump connection manner to obtain a voice mapping feature.
[0075] In the embodiment of the present application, the jump connection means directly adding the output to the input of a certain layer of the neural network, which is manifested as: in the forward propagation process of the network, the output of a certain layer skips some neural network layers in the neural network module, and the output is used as the input of the following neural network layer. For example, the output of the first neural network layer is used as the input of the third neural network layer. This connection is usually implemented through an addition operation, adding the input and the output. For example, the original input feature of the third neural network layer is the output feature of the second neural network layer. After adding the output feature of the first neural network layer to the original input feature of the third neural network layer through the jump connection method, the final input feature of the third neural network layer is obtained.
[0076] Based on the speech feature vectors after multiple feature transformations, by using a neural network module to perform feature mapping on the speech feature vectors, during the mapping process, skip connections are used to connect the neural network layers included in the neural network module to achieve feature transmission. Usually, traditional neural network modules increase the depth of the network by stacking neural network layers, thereby improving the noise reduction effect of the module. However, when the number of network layers increases to a certain amount, since the neural network is performing backpropagation, which requires continuous propagation of gradients, when the network depth increases, the gradients will gradually disappear, resulting in the inability to adjust the weights of the previous network layers, leading to a reduction in the noise reduction effect of the module. By using skip connections to perform feature mapping on the speech feature vectors, the problem of gradient disappearance can be solved and the network learning process can be accelerated.
[0077] In some embodiments, a neural network module with temporal characteristics (such as GRU, etc.) can also be used. A neural network module with temporal characteristics can capture long-term dependencies in sequential data, and speech feature vectors belong to sequential data. By using a neural network module with temporal characteristics, the correlation between speech feature vectors can be effectively learned, improving the learning performance of the neural network module.
[0078] Step S104, perform feature fusion on the speech feature vectors and the speech mapping features to obtain the gain information of the speech signal to be enhanced.
[0079] In the embodiments of the present application, the gain information is represented as a scaling factor between the enhanced speech signal and the speech signal to be enhanced. The gain information can be the ratio of the energy spectrum of the enhanced speech signal to the energy spectrum of the speech signal to be enhanced, or the ratio of the energy of the enhanced speech signal to the energy of the speech signal to be enhanced. It can be a gain array composed of any values between 0 and 1, including multiple gain values. For each speech signal to be enhanced, there corresponds a different gain value, that is, the dimension of the gain information is the same as the feature dimension of each speech signal to be enhanced. By using a feature fusion module to perform feature fusion on the speech feature vectors and the speech mapping features, the gain information of the speech signal to be enhanced is obtained, that is, the gain value corresponding to each speech signal to be enhanced is obtained.
[0080] Feature fusion can be to jointly model the speech feature vectors after feature transformation and the speech mapping features after feature mapping by using a neural network module. By capturing different feature information in the speech feature vectors and the speech mapping features, it helps the model better understand the data, thereby improving the performance and generalization ability of the model.
[0081] In some embodiments, signal enhancement can also be performed on the speech signal to be enhanced based on the gain information to obtain an enhanced speech signal of the speech signal to be enhanced. Based on the power of the speech signal to be enhanced at each frequency component, the power spectrum of the speech signal to be enhanced is obtained, and signal enhancement is achieved by multiplying the gain information by the power spectrum of the speech signal to be enhanced. In the frequency domain, each gain value in the gain information is also a gain value at different frequency components, and this gain value corresponds to the power of the speech feature to be enhanced at each frequency component. Multiply the gain value of the speech feature to be enhanced at each frequency component by the power to obtain an enhanced audio signal. In some embodiments, the enhanced audio signal can be converted from the frequency domain to the time domain through an inverse fast Fourier transform to obtain an enhanced speech signal of the speech feature to be enhanced, thereby achieving signal enhancement of the speech feature to be enhanced.
[0082] In some embodiments, if band compression is performed when generating the speech feature to be enhanced, then correspondingly, before performing signal enhancement on the speech signal to be enhanced, band decompression processing needs to be performed on the gain information. For example, when generating the speech feature to be enhanced, the 513-dimensional feature can be compressed to 128 dimensions through band compression to obtain a 128-dimensional speech feature to be enhanced, and the gain information output after the network model processes is also 128-dimensional. Therefore, before performing signal enhancement on the speech signal to be enhanced, the 128-dimensional gain information needs to be decompressed to 513 dimensions and then multiplied by the power spectrum of the speech feature to be enhanced before band compression to obtain an enhanced audio signal.
[0083] The speech signal processing method provided by the embodiments of the present application first extracts features from the speech signal to be enhanced, and circularly performs multiple feature transformations on the speech feature to be enhanced obtained after feature extraction; then uses a skip connection method to perform feature mapping on the speech feature vector obtained by feature transformation; and fuses the speech feature vector and the speech mapping feature obtained by feature mapping to obtain the gain information of the speech signal to be enhanced; finally, an enhanced speech signal of the speech feature to be enhanced is obtained through the signal enhancement process. In this way, by circularly performing multiple feature transformations, the network depth and the number of parameters are increased, and the speech feature to be enhanced is learned more effectively; through the skip connection in the feature mapping process, the continuity and correlation between speech feature vectors are ensured; the feature transformation result and the feature mapping result are fused to obtain accurate gain information, ensuring the enhancement effect on the speech signal to be enhanced during the signal enhancement process and improving the speech signal quality.
[0084] The following will describe the voice signal processing method in the embodiments of the present application in combination with the interaction between the terminal and the server in the voice signal processing system. It should be noted that the voice signal processing method here is a voice signal processing method implemented through the interaction between the terminal and the server, which is substantially the same as the voice signal processing method executed by the server in the above embodiments. The only difference is that the embodiments of the present application also describe the actions performed by the terminal during the execution of the voice signal processing method. Moreover, some steps can be executed either by the terminal or by the server. Therefore, for the steps that are the same in content but have different execution entities in this embodiment and the above embodiments, this embodiment is only an exemplary illustration, and in the implementation process, it can be executed by any one of the execution entities, and the embodiments of the present application do not make any limitations in this regard.
[0085] Figure 4 is another optional flowchart of the voice signal processing method provided by the embodiments of the present application. As Figure 4 shown, the method includes the following steps S201 to step S216:
[0086] Step S201, the terminal receives a voice signal processing operation input by the user.
[0087] In the embodiments of the present application, the user can input a voice signal processing operation on the client side of the voice signal processing application. In the voice signal processing application, voice signal processing functions can be provided. The user (who can be a voice processing personnel or a voice signal processing designer) can input a voice signal processing operation on the voice signal processing function page to trigger a voice signal processing request.
[0088] In some embodiments, when the user inputs a voice signal processing operation, the user can also input the voice signal to be enhanced at the same time. Or, in other embodiments, correspondingly when the user inputs a voice signal processing operation, the user can also input the feature data of the voice signal to be enhanced after any one or all of the feature preprocessings. The feature preprocessing can include frequency domain transformation, audio feature extraction, and frequency band compression.
[0089] Step S202, the terminal generates a voice signal processing request in response to the voice signal processing operation.
[0090] In the embodiments of the present application, the data input by the user can be encapsulated into the voice signal processing request. For example, the voice signal to be enhanced input by the user can be encapsulated into the voice signal processing request, or the feature data of the voice signal to be enhanced after any one or all of the feature preprocessings input by the user can be encapsulated into the voice signal processing request.
[0091] Step S203, the terminal sends the voice signal processing request to the server.
[0092] Step S204: In response to the voice signal processing request, the server performs a frequency-domain transformation on the voice signal to be enhanced to obtain voice audio features.
[0093] In the embodiments of the present application, if the voice signal to be enhanced is encapsulated in the voice signal processing request, the voice signal to be enhanced can be parsed and a frequency-domain transformation can be performed on the voice signal to be enhanced, that is, the voice signal to be enhanced is converted from the time domain to the frequency domain through FFT to obtain the frequency-domain features of the voice signal to be enhanced, that is, the voice audio features; if the result of the frequency-domain transformation of the voice signal to be enhanced is encapsulated in the voice signal processing request, the voice audio features can be directly parsed. Among them, the voice signal to be enhanced can be an ultra-wideband voice signal or an ordinary voice signal.
[0094] For example, a 1024-point FFT can be used to perform a frequency-domain transformation on the voice signal to be enhanced to obtain voice audio features. The real part of the voice audio features is 513 features, and the imaginary part is also 513 features, for a total of 513 * 2 features.
[0095] In the embodiments of the present application, by performing a frequency-domain transformation on the voice signal to be enhanced, the voice audio features in the frequency domain are obtained, converting the complex time-domain signal into a frequency-domain signal that is easy to analyze, and obtaining the voice audio features of multiple frequency components. This feature includes the real part and the imaginary part of the amplitude. The real part represents the amplitude of the voice signal to be enhanced, and the imaginary part represents the change of the phase of the voice signal to be enhanced with frequency, which facilitates subsequent feature analysis of the voice signal to be enhanced.
[0096] Step S205: The server extracts audio features from the voice audio features to obtain an audio feature vector.
[0097] In the embodiments of the present application, for the voice audio features, in order to learn a feature representation with good generalization ability, a neural network with good feature extraction ability can be used to extract features from the voice audio features. For example, a CNN. Or, if the result of the voice signal to be enhanced after passing through frequency-domain transformation and audio feature extraction in sequence is encapsulated in the voice signal processing request, the audio feature vector can be directly parsed.
[0098] Step S206: The server performs frequency band compression on the audio feature vector to obtain the voice feature to be enhanced.
[0099] In the embodiments of the present application, band compression refers to performing dimensionality reduction processing on an audio feature vector, and performing band compression on the audio feature vector through a band compression layer. For example, the band compression layer may include a band-pass filter, and the band compression layer may divide each frequency of the audio feature vector into a set number of frequency bands, so as to obtain an enhanced speech feature including the set number of frequency bands. Alternatively, if the enhanced speech signal encapsulated in the speech signal processing request is the result after frequency domain transformation, audio feature extraction, and band compression in sequence, the enhanced speech feature can be directly parsed.
[0100] For example, when the dimension of the audio feature vector is 513 dimensions and there are 513 frequency points, the set number of frequency bands may be less than 513, for example, 128. The band compression layer may be an Equivalent Rectangular Bandwidth (ERB), and the ERB is used to divide each frequency of the audio feature vector into 128 frequency bands, reducing the feature dimension from 513 dimensions to 128 dimensions. Through the ERB, the 513-dimensional audio feature vector can be converted into a 128-dimensional enhanced speech feature.
[0101] Through the band compression layer, the audio feature vector can be dimensionally reduced to obtain the enhanced speech feature after dimensionality reduction, thereby realizing feature dimensionality reduction. In this way, by calculating with the features after dimensionality reduction in subsequent steps, the computational complexity of the model can be reduced, and the model operation speed can be improved.
[0102] Step S207, the server repeatedly performs a first feature transformation on the enhanced speech feature through a first feature transformation module with a specific number of channels to obtain a first speech feature vector.
[0103] In the embodiments of the present application, during multiple first feature transformations, the specific number of channels of the first feature transformation module increases successively. The specific number of channels is the number of output channels of the convolutional layer in the first feature transformation module, that is, the number of convolutional kernels in the convolutional layer.
[0104] In some embodiments, the first feature transformation module includes a first convolutional layer, a first normalization layer, a first activation layer, and a first attention mechanism layer; see Figure 5 , Figure 5 shows that in step S207, the server repeatedly performs a first feature transformation on the enhanced speech feature through a first feature transformation module with a specific number of channels to obtain a first speech feature vector, which can be implemented through the following steps S2071 to S2074:
[0105] Step S2071, during each first feature transformation, perform convolutional processing on the enhanced speech feature through the first convolutional layer to obtain a first convolutional feature.
[0106] In the embodiment of the present application, the first convolutional layer includes a two-dimensional convolutional layer, and the number of two-dimensional convolutional layers can be multiple layers, and the number of convolutional layers is not limited herein. During each first feature transformation, the first convolutional layer is used to perform convolutional processing on the speech feature to be enhanced, that is, to extract features from the speech feature to be enhanced, so as to obtain the first convolutional feature after feature extraction, which is convenient for improving the learning effect of the subsequent model.
[0107] Step S2072: Normalize the first convolutional feature through the first normalization layer to obtain the first normalized feature.
[0108] In the embodiment of the present application, the role of the first normalization layer is to normalize the first convolutional feature. The first normalization layer can have preset learning parameters. By using the normalization layer with preset learning parameters to normalize the first convolutional feature, the first convolutional feature after normalization is obtained, that is, the first normalized feature. The first normalized feature can make the distribution of the input of each subsequent neural network layer consistent during the feature learning process, that is, the distribution of the input data of each layer in the model is relatively stable, and the model learning speed is accelerated.
[0109] Step S2073: Activate the first normalized feature through the first activation layer to obtain the first activation feature.
[0110] In the embodiment of the present application, the first activation layer is used to process the linear output of the previous layer through a non-linear activation function to simulate any function. For example, the activation function in the activation layer can include PReLU function, ReLU function, tanh function or sigmoid function, etc., and the activation function can be located between at least two convolutional layers. By activating the first normalized feature through the first activation layer, the first normalized feature after activation processing is obtained, that is, the first activation feature, thereby enhancing the representation ability of the model.
[0111] Step S2074: Allocate features to the first activation feature through the first attention mechanism layer to obtain the first speech feature vector.
[0112] In the embodiment of the present application, the first attention mechanism layer is used to adjust the importance degree of the feature values in the first activation feature through learnable weights, highlight the important features in the first activation feature, and suppress the irrelevant features in the first activation feature. The first attention mechanism layer can combine a traditional neural network with an attention mechanism. For example, a convolutional neural network and an attention mechanism are combined to form the first attention mechanism layer, and an attention module can be introduced between convolutional layers to automatically learn the important features in the channels according to the content of the first activation feature and make more full use of these features, thereby improving the performance of the model.
[0113] In some embodiments, the number of times of the first feature transformation can be set according to the actual model performance, that is, the first feature transformation process can be cycled multiple times. For example, if the number of loop iterations is set to 3, the output result of the first attention mechanism layer is input into the above several network layers again, and so on, until after 3 loop iterations, the first speech feature vector is finally output. And, the number of output channels of the convolutional layer can be changed during each loop iteration. For example, the set number of output channels can be incremented successively, and can be 32, 64, and 128 respectively. By looping the first feature transformation multiple times and setting specific channel numbers for each first feature transformation, the network depth and the number of parameters can be increased, and the speech features to be enhanced can be learned more effectively.
[0114] Step S208, the server performs at least one attention feature transformation on the first speech feature vector based on the attention mechanism through the attention module to obtain a second speech feature vector.
[0115] In the embodiments of the present application, the attention module may include at least one attention mechanism layer, and performs at least one attention feature transformation on the first speech feature vector based on the attention mechanism. The attention mechanism layer can also be a combination of a traditional neural network and the attention mechanism. The number of times of the attention feature transformation is set according to the actual model performance. For example, it can be set to 3 times, that is, the first speech feature vector is sequentially subjected to attention feature transformation through three attention mechanism layers. During the multiple attention feature transformation processes, the number of output channels of the attention mechanism layer can remain unchanged. For example, the number of output channels is maintained at 128, or it can also be changed according to requirements, which is not limited herein. The model can further focus on the correlation between different channels, and finally combine these features to form the final output result, that is, obtain the second speech feature vector.
[0116] Step S209, the server performs multiple second feature transformations on the second speech feature vector in a loop through the second feature transformation module with a specific number of channels to obtain a speech feature vector.
[0117] In the embodiments of the present application, during the multiple second feature transformations, the specific number of channels of the second feature transformation module decreases successively. The specific number of channels is the number of output channels of the convolutional layer in the second feature transformation module, that is, the number of convolutional kernels in the convolutional layer.
[0118] In some embodiments, the specific number of channels of the first feature transformation module when performing the last first feature transformation on the speech features to be enhanced is the same as the specific number of channels of the second feature transformation module when performing the first second feature transformation on the second speech feature vector. Moreover, the dimension of the speech feature vector obtained after performing the last second feature transformation on the second speech feature vector is the same as the dimension of the speech features to be enhanced. The second feature transformation module includes a second convolutional layer, a second normalization layer, a second activation layer, and a second attention mechanism layer; see Figure 6 , Figure 6 shows that in step S209, the server cyclically performs multiple second feature transformations on the second speech feature vector through the second feature transformation module with a specific number of channels to obtain a speech feature vector, which can be implemented through the following steps S2091 to S2094:
[0119] Step S2091: At each second feature transformation, perform convolutional processing on the second speech feature vector through the second convolutional layer to obtain a second convolutional feature.
[0120] Step S2092: Through the second normalization layer, perform normalization processing on the second convolutional feature to obtain a second normalized feature.
[0121] Step S2093: Through the second activation layer, perform activation processing on the second normalized feature to obtain a second activation feature.
[0122] Step S2094: Through the second attention mechanism layer, perform feature allocation on the second activation feature to obtain a speech feature vector.
[0123] In the embodiments of the present application, the structures of the second convolutional layer, the second normalization layer, the second activation layer, and the second attention mechanism layer in the second feature transformation module are similar to those of the first convolutional layer, the first normalization layer, the first activation layer, and the first attention mechanism layer in the corresponding first feature transformation module. Therefore, the content of each network layer in the second feature transformation module will not be elaborated in detail. The number of times of the second feature transformation can also be set according to the actual model performance, that is, the process of the second feature transformation can be cycled multiple times. For example, if the number of loop iterations is set to 3 times, the output result of the second attention mechanism layer is input into the above-mentioned several network layers again, and so on, until after 3 loop iterations, the speech feature vector is finally output. Moreover, the number of output channels of the convolutional layer can also be changed during each loop iteration. For example, the set number of output channels can be gradually decreased, and can be 128, 64, and 32 respectively. By cyclically performing the second feature transformation multiple times and setting a specific number of channels for each second feature transformation to reduce the channel dimension, the feature dimension of the speech feature vector is kept consistent with the input feature of the first feature transformation module (i.e., the speech features to be enhanced).
[0124] Step S210: The server performs feature mapping on the speech feature vector in a skip connection manner through a multi-layer gated recurrent unit with a skip connection relationship to obtain a speech mapping feature.
[0125] In the embodiment of the present application, the multi-layer gated recurrent unit with a skip connection relationship has N layers; the input feature of the first-layer gated recurrent unit is the speech feature vector; the input feature of the Nth-layer gated recurrent unit is the output features of all the gated recurrent units connected before the Nth-layer gated recurrent unit and the speech feature vector; N is a positive integer greater than 1.
[0126] In some embodiments, according to the final enhanced speech processing effect, the number of layers of the multi-layer gated recurrent unit can be changed. For example, referring to Figure 7A , Figure 7A shows a schematic diagram of the implementation process of feature mapping of the three-layer gated recurrent unit provided by the embodiment of the present application. Among them, the number of layers of the gated recurrent unit is set to 3 layers, namely gated recurrent unit 701, gated recurrent unit 702, and gated recurrent unit 703. The input feature of the first-layer gated recurrent unit 701 is the speech feature vector. Based on the skip connection, the speech feature vector and the output feature of the first-layer gated recurrent unit 701 are jointly used as the input feature of the second-layer gated recurrent unit 702, and the speech feature vector, the output feature of the first-layer gated recurrent unit 701, and the output feature of the second-layer gated recurrent unit 702 are jointly used as the input feature of the third-layer gated recurrent unit 703. Referring to Figure 7B , Figure 7B shows a schematic diagram of the implementation process of feature mapping of the four-layer gated recurrent unit provided by the embodiment of the present application. Among them, the number of layers of the gated recurrent unit is set to 4 layers, namely gated recurrent unit 701, gated recurrent unit 702, gated recurrent unit 703, and gated recurrent unit 704. For the newly added fourth-layer gated recurrent unit 704, the speech feature vector, the output feature of the first-layer gated recurrent unit 701, the output feature of the second-layer gated recurrent unit 702, and the output feature of the third-layer gated recurrent unit 703 are jointly used as the input feature of the fourth-layer gated recurrent unit 704. Referring to Figure 7C , Figure 7C shows a schematic diagram of the implementation process of feature mapping of the five-layer gated recurrent unit provided by the embodiment of the present application. Among them, the number of layers of the gated recurrent unit is set to 5 layers, namely gated recurrent unit 701, gated recurrent unit 702, gated recurrent unit 703, gated recurrent unit 704, and gated recurrent unit 705. For the newly added fifth-layer gated recurrent unit 705, the speech feature vector, the output feature of the first-layer gated recurrent unit 701, the output feature of the second-layer gated recurrent unit 702, the output feature of the third-layer gated recurrent unit 703, and the output feature of the fourth-layer gated recurrent unit 704 are jointly used as the input feature of the fifth-layer gated recurrent unit 705.
[0127] In some embodiments, for the Nth layer gated recurrent unit, first, obtain the output features of all the gated recurrent units connected before the Nth layer gated recurrent unit and the speech feature vector; perform a vector summation operation on the output features of all the gated recurrent units connected before the Nth layer gated recurrent unit and the speech feature vector to obtain the input features of the Nth layer gated recurrent unit; perform feature mapping on the input features through the Nth layer gated recurrent unit to obtain the output features of the Nth layer gated recurrent unit.
[0128] For example, for the third layer gated recurrent unit 703, it is necessary to obtain the output features of all the gated recurrent units connected before the third layer gated recurrent unit, that is, the output features of the first layer gated recurrent unit 701, the output features of the second layer gated recurrent unit 702, and the speech feature vector, and perform vector addition on the output features of the first layer gated recurrent unit 701, the output features of the second layer gated recurrent unit 702, and the speech feature vector, and use the result of the vector addition as the input features of the third layer gated recurrent unit 703. Perform feature mapping on the input features through the third layer gated recurrent unit 703 to obtain the output features of the third layer gated recurrent unit 703.
[0129] In some embodiments, the number of neuron nodes of the multi-layer gated recurrent unit can also be changed so that the number of neuron nodes of each layer of the gated recurrent unit is different. For example, the number of neuron nodes of each layer of the gated recurrent unit can increase successively. When the feature dimension of the output features of the last layer of the gated recurrent unit with skip connection is higher than the feature dimension threshold, a gated recurrent unit can be added after the last layer of the gated recurrent unit to perform dimensionality reduction processing on the output features to obtain the final speech mapping features. The feature dimension threshold can be set according to actual needs and is not limited here.
[0130] The feature mapping process is implemented through the gated recurrent unit. The flow of feature information can be controlled through learnable gates, and the feature information at the current moment in the speech feature vector can be associated with the feature information at the previous and subsequent moments, better capturing the dependencies with a large time step distance in the time series, thereby ensuring the continuity and correlation between the speech feature vectors.
[0131] Step S211, the server performs feature fusion on the speech feature vector and the speech mapping feature to obtain the gain information of the speech signal to be enhanced.
[0132] In the embodiments of the present application, in the feature fusion process, the speech feature vectors obtained after multiple feature transformations and the speech mapping features obtained after feature mapping can be merged by concatenation or addition, and then the result after feature merging is processed by a feature fusion layer. By fusing different feature information in the speech feature vectors and the speech mapping features, the result after the fusion process is obtained, that is, the gain information of the speech signal to be enhanced. The feature fusion layer can be different neural network layers. For example, a two-dimensional convolutional layer. The feature fusion process helps the neural network model better understand the data, thereby improving the performance and generalization ability of the model.
[0133] Step S212: The server decompresses the gain information in the frequency band to obtain the decompressed gain information.
[0134] In the embodiments of the present application, the frequency band decompression is the inverse process of frequency band compression. Since in the above process, in order to reduce the computational complexity of the model, the frequency band compression layer is used to transform the high-dimensional audio feature vector into a low-dimensional speech feature to be enhanced as the input feature of the model. Therefore, in order to obtain the enhanced speech signal corresponding to the speech signal to be enhanced subsequently, it is necessary to decompress the gain information through the frequency band decompression layer, and restore the gain information to a high dimension, that is, expand from a low dimension to a high dimension. For example, ERB can be used as the frequency band decompression layer, and ERB is used to expand the 128-dimensional gain information to 513-dimensional gain information.
[0135] Step S213: The server performs a multiplication operation on the decompressed gain information and the power spectrum of the speech audio feature to obtain the enhanced audio feature.
[0136] In the embodiments of the present application, the decompressed gain information includes multiple gain values, and the power value of the speech audio feature at each frequency component is obtained. There is a preset mapping relationship between the power value of the speech audio feature at each frequency component and one gain value in the decompressed gain information, that is, each power value corresponds to one gain value. Based on the preset mapping relationship, a multiplication operation is performed on each gain value and each power value to obtain the enhanced audio feature.
[0137] Step S214: The server performs an inverse frequency domain transformation on the enhanced audio feature to obtain the enhanced speech signal of the speech feature to be enhanced.
[0138] In the embodiments of the present application, the inverse frequency domain transformation and the frequency domain transformation are inverse processes of each other. The inverse frequency domain transformation is to transform the enhanced audio feature in the frequency domain from the frequency domain to the time domain to obtain the enhanced speech signal in the time domain.
[0139] Step S215: The server sends the enhanced speech signal of the speech feature to be enhanced to the terminal.
[0140] Step S216: The terminal outputs the enhanced speech signal of the speech feature to be enhanced.
[0141] In the embodiments of the present application, when performing speech enhancement through a speech signal processing model, first, an audio feature vector of the speech signal to be enhanced is extracted. Then, in order to reduce the subsequent computational complexity, the audio feature vector is subjected to band compression to obtain the speech feature to be enhanced with reduced dimensions. Further, through the first feature transformation module, the attention module, and the second feature transformation module of the speech enhancement model, the speech feature to be enhanced is subjected to feature transformation in a cyclic iterative manner to obtain a speech feature vector, increasing the depth of feature transformation without increasing the model structure to ensure the model processing effect; the speech feature vector is subjected to feature mapping in a jump connection manner to obtain a speech mapping feature, ensuring the continuity and correlation between speech feature vectors; finally, based on the speech mapping feature and the speech feature vector, estimated gain information is obtained to perform speech enhancement on the speech signal to be enhanced. Therefore, the embodiments of the present application can reduce the computational complexity while ensuring the processing effect, so as to improve the computational speed, thereby meeting the requirements of real-time computing.
[0142] In some embodiments, the above speech signal processing method can be implemented through a speech signal processing model. Refer to Figure 8 , Figure 8 which shows a schematic flowchart of the speech signal processing model training method provided by the embodiments of the present application. The speech signal processing model training method of the embodiments of the present application can be executed by a model training module. Among them, the model training module can be a module in an electronic device for implementing the speech signal processing method, that is, the speech signal processing model training method can be executed by a terminal or a server; of course, the model training module can also be a module in other electronic devices different from the electronic device for implementing the speech signal processing method, that is, the speech signal processing model training method can be executed by other terminals or other servers. As Figure 8 shown, the speech signal processing model is trained through the following steps S301 to S307:
[0143] Step S301, obtain sample data.
[0144] In the embodiments of the present application, the sample data includes a first type of speech sample, a second type of speech sample corresponding to the first type of speech sample, and the true gain between the first type of speech sample and the second type of speech sample; the first type of speech sample is a clean speech sample without noise, and the second type of speech sample is a speech sample obtained by adding noise to the first type of speech sample. The first type of speech sample and the first type of speech sample may be ultra-wideband speech signals or ordinary speech signals. The true gain between the first type of speech sample and the second type of speech sample can be obtained by dividing the energy spectrum of the first type of speech sample by the energy spectrum of the second type of speech sample signal, or it can also be obtained by dividing the energy of the first type of speech sample by the energy of the second type of speech sample signal, that is, the true gain between the first type of speech sample and the second type of speech sample is the ratio of the energy spectra of the two or the ratio of the energies of the two, and the numerical range is between 0 and 1.
[0145] Step S302: Perform feature preprocessing on the first type of speech sample and the second type of speech sample to obtain the speech features to be enhanced for the sample, and input the speech features to be enhanced for the sample into the speech signal processing model.
[0146] In the embodiments of the present application, audio feature extraction is respectively performed on the first type of speech sample and the second type of speech sample to correspondingly obtain a first type of sample audio feature vector and a second type of sample audio feature vector. The audio feature extraction process can be implemented through an audio feature extraction layer, that is, inputting the first type of speech sample and the second type of speech sample into the same audio feature extraction layer for audio feature extraction, or the first type of speech sample and the second type of speech sample can also be input into different audio feature extraction layers for audio feature extraction. For example, the audio feature extraction layer can be a convolutional layer. Feature splicing is performed on the first type of sample audio feature vector and the second type of sample audio feature vector to obtain the spliced speech features for the sample, that is, the first type of sample audio feature vector and the second type of sample audio feature vector are spliced together in sequence, and then the feature data is integrated through a feature splicing layer to obtain the spliced speech features for the sample. For example, the feature splicing layer can be a convolutional layer. Band compression is performed on the spliced speech features for the sample to obtain the speech features to be enhanced for the sample, that is, the high-dimensional spliced speech features for the sample are transformed into low-dimensional speech features to be enhanced for the sample through a band compression layer. Finally, the low-dimensional speech features to be enhanced for the sample are input into the speech signal processing model to reduce the dimensionality of the input features of the model and reduce the computational complexity of the model.
[0147] Step S303: Through the feature transformation network of the speech signal processing model, perform multiple feature transformations on the speech features to be enhanced for the sample in a loop to obtain the sample speech feature vector.
[0148] In the embodiments of the present application, the feature transformation network of the voice signal processing model includes a first feature transformation module, an attention module, and a second feature transformation module. Both the first feature transformation module and the second feature transformation module have specific numbers of channels, and the numbers of channels of the two modules can change during multiple feature transformation processes. For example, the number of channels of the first feature transformation module can increase successively during the loop iteration process, and the number of channels of the second feature transformation module can decrease successively during the loop iteration process. The number of channels of the attention module can also be changed or remain unchanged according to actual requirements.
[0149] In some embodiments, through the first feature transformation module with a specific number of channels, multiple first feature transformations are performed on the sample voice features to be enhanced in a loop to obtain the sample first voice feature vector; wherein, during the multiple first feature transformations, the specific number of channels of the first feature transformation module increases successively; through the attention module, at least one attention feature transformation is performed on the sample first voice feature vector based on the attention mechanism to obtain the sample second voice feature vector; through the second feature transformation module with a specific number of channels, multiple second feature transformations are performed on the sample second voice feature vector in a loop to obtain the sample voice feature vector; wherein, during the multiple second feature transformations, the specific number of channels of the second feature transformation module decreases successively.
[0150] Step S304: Through the feature mapping network of the voice signal processing model, the sample voice feature vector is feature-mapped in a skip connection manner to obtain the sample voice mapped feature.
[0151] In the embodiments of the present application, the feature mapping network of the voice signal processing model includes multiple gated recurrent units with skip connection relationships, and the gated recurrent units are connected between the input features and the output features in a skip connection manner. Through the multiple gated recurrent units with skip connection relationships, the sample voice feature vector is feature-mapped in a skip connection manner to obtain the sample voice mapped feature; wherein, the multiple gated recurrent units with skip connection relationships have N layers; the input feature of the first layer of gated recurrent unit is the sample voice feature vector; the input feature of the Nth layer of gated recurrent unit is the output features of all the gated recurrent units connected before the Nth layer of gated recurrent unit and the sample voice feature vector; N is a positive integer greater than 1. Through this feature mapping process, the problem of gradient disappearance during the model training process is solved and the model training process is accelerated.
[0152] Step S305: Through the feature fusion network of the voice signal processing model, the sample voice feature vector and the sample voice mapped feature are feature-fused to obtain the sample estimated gain of the sample voice signal to be enhanced.
[0153] In the embodiments of the present application, the feature fusion network of the voice signal processing model includes a feature fusion layer. The number of feature fusion layers can be a single layer or multiple layers. Through the feature fusion layer, the sample voice feature vector and the sample voice mapping feature are spliced and fused, and different feature information in the sample voice feature vector and the sample voice mapping feature is fused to obtain a fused result, that is, the sample estimated gain of the sample voice signal to be enhanced. Here, the network type of the feature fusion layer is not limited. Through the feature fusion network of the voice signal processing model, an accurate sample estimated gain can be obtained, thereby improving the signal quality of the enhanced voice signal.
[0154] Step S306: Input the sample estimated gain and the true gain into the loss model for loss calculation to obtain a loss result.
[0155] In the embodiments of the present application, based on the sample estimated gain and the true gain, a cost function of the loss model is constructed and loss calculation is performed to obtain a loss result. For example, the cost function can be a Mean Squared Error (MSE) function. The loss result is used to measure the degree of inconsistency between the sample estimated gain and the true gain of the model, that is, to calculate the gap between the forward calculation result (i.e., the sample estimated gain) of each iteration of the model and the true value (i.e., the true gain), so as to guide the next training in the correct direction.
[0156] Step S307: Update the model parameters in the voice signal processing model based on the loss result to obtain a trained voice signal processing model.
[0157] In the embodiments of the present application, according to the derivative of the cost function, the loss result is backpropagated along the direction of the minimum gradient to update the model parameters in the feature transformation network, the feature mapping network, and the feature fusion network, such as the respective weight values in the feature transformation network, the feature mapping network, and the feature fusion network. A loss result threshold is preset. When the loss result is less than the preset loss result threshold, the iterative training is stopped, that is, the model parameter update is stopped; alternatively, a maximum iteration number threshold can be preset. When the iteration number exceeds the maximum iteration number threshold, the model parameter update is stopped to obtain a trained voice signal processing model.
[0158] In the embodiments of the present application, the preprocessed sample speech features to be enhanced are input into a speech signal processing model. Through three processing networks in the speech signal processing model, namely, a feature transformation network, a feature mapping network, and a feature fusion network, the sample estimated gain of the sample speech features to be enhanced is obtained. Based on the loss calculation between the sample estimated gain and the true gain, an accurate speech signal processing model is trained, so as to use the speech signal processing model to obtain the optimal sample estimated gain, and then based on the sample estimated gain, the effective enhancement of the sample speech signal to be enhanced is realized, and the noise reduction effect of the speech signal processing model is improved.
[0159] Next, the exemplary application of the embodiments of the present application in a practical application scenario will be described.
[0160] The embodiments of the present application provide a speech signal processing method, which involves an ultra-wideband speech signal processing model with serial fusion of a CNN and a GRU. The speech signal processing model can be applied to real-time communication conference software. When the call environment is relatively noisy, using the speech signal processing model can effectively remove the background environmental noise of the voice call, thereby ensuring the quality and intelligibility of the voice call. When the microphone is turned on, the sound signal is collected through the microphone, and then the collected sound signal is input into the speech signal processing model. After the speech signal processing model performs enhancement processing on the sound signal, a relatively "clean" speech signal can be obtained. This signal is regarded as the enhanced speech signal, effectively retaining the voice of the main speaker and removing the redundant background noise.
[0161] Figure 9 For the neural network model training structure proposed in the embodiments of the present application, for ultra-wideband speech signals (sampling rate of 32 KHz and bandwidth of 16 KHz). The input feature data for network training is a set (a pair), which are clean features (i.e., the first type of speech samples mentioned above) and noisy features (i.e., the second type of speech samples mentioned above). The real and imaginary parts of the clean features and the noisy features are obtained respectively by using a 1024-point FFT operation (i.e., the above-mentioned frequency domain transformation). The real part of the clean features and the noisy features has 513 features, and the imaginary part also has 513 features. Therefore, the total number of input feature data for network training is 513 * 4 features.
[0162] The specific implementation steps of network training are as follows: Input training features: 513 * 2 clean features and 513 * 2 noisy features. The clean features and the noisy features are respectively subjected to feature extraction (i.e., the above-mentioned audio feature extraction) through a CNN module (i.e., a convolutional neural network), and the feature extraction results are concatenated together in sequence. Then, the concatenated feature data is fused through a two-dimensional CNN layer (i.e., the above-mentioned feature concatenation). The integrated feature data is passed through a frequency band compression module, i.e., an ERB module. The ERB divides 128 frequency bands, reducing the feature dimension of the fused and concatenated feature data from 513 dimensions to 128 dimensions (i.e., the above-mentioned frequency band compression), obtaining the ERB result. This module further reduces the network complexity and simultaneously compresses the frequency band and reduces the feature dimension. The ERB result is input into module A (i.e., the above-mentioned first feature transformation module), and passes through a series of sub-modules in module A: CNN, batch normalization, activation function (PReLU function), CNN, and Attention. This process is looped three times in total. The number of channels of the CNN increases sequentially each time, being 32, 64, and 128 respectively, aiming to increase the network depth and the number of parameters, so that the signal features can be learned more effectively.
[0163] The result after three loops of module A is input into three groups of Attention modules (i.e., the above-mentioned attention modules, Figure 9 the attention mechanism modules in it). The technology of the Attention module is completed using two-dimensional CNN. The number of channels of each group of Attention modules remains unchanged, being 128, 128, and 128 respectively. Then, the results output by the three groups of Attention modules are input into module B (i.e., the above-mentioned second feature transformation module). Similarly, the network architecture in module B is similar to that in module A, and it is looped three times, but the number of channels of the CNN decreases sequentially each time, being 128, 64, and 32 respectively. The function is to reduce the channel dimension and keep it consistent with the input features of module A.
[0164] The output features of module B are used as the input features of the GRU module and passed into Figure 9 the GRU module on the right in it (i.e., the above-mentioned feature mapping network). The GRU module corresponds to Figure 9 module C in it. The GRU module includes four layers of GRU units (i.e., gated recurrent units) (the number of nodes is 70, 100, 180, and 128 respectively). The first three layers of GRU units are connected using skip connections to ensure temporal continuity and correlation. The fourth layer of GRU unit 901 can be used to perform feature dimension reduction on the features output after skip connection, reducing the feature dimension of the output features of the GRU module.
[0165] Finally, using the feature fusion network (i.e., the above-mentioned feature fusion network), the two-dimensional CNN module fuses the output features of the GRU module and the output features of module B, and outputs the estimated 128-dimensional gain information, that is, the ERB gain (i.e., the above-mentioned gain information), and obtains the 513-dimensional ERB gain (i.e., the above-mentioned decompressed gain information) through ERB expansion. Compare the 513-dimensional ERB gain with the corresponding 513-dimensional true gain (which can be calculated from the clean speech and the noisy speech signal, that is, the true gain between the above-mentioned first type of speech sample and the second type of speech sample) using the MSE cost function, that is, construct the MSE cost function using the 513-dimensional ERB gain and the true gain, and optimize the overall network architecture to obtain the model parameters of the entire network.
[0166] After the network model parameters are obtained through the above network training process, the trained network model can be used for online enhancement processing. Only the corresponding real and imaginary parts of the features of 513*2 (consistent with the above features) need to be extracted from the noisy speech signal (i.e., the above-mentioned speech signal to be enhanced). Then, input the features into the trained network model, and the network will perform adaptive parameter adaptation to obtain the optimal gain information. Finally, multiply the gain by the power spectrum of the noisy signal, and then use the inverse FFT to obtain the finally enhanced speech signal (i.e., the above-mentioned enhanced speech signal).
[0167] Figure 10 is the original speech signal. After processing the original speech signal through the speech signal processing process of the embodiment of the present application, the obtained enhanced speech signal is as Figure 11 shown. Combining Figure 10 and Figure 11 the comparison of the enhancement effects, it can be seen that the noise signal in the original speech signal is removed from the enhanced speech signal, and the enhancement effect is obvious. In the speech segment, the noise removed by the algorithm of the embodiment of the present application is relatively clean. At the same time, the number of model parameters is about 930,000, which is actually a low-complexity enhancement algorithm.
[0168] In the embodiment of the present application, by combining the strong feature extraction ability of CNN and the temporal characteristics of GRU, and introducing the attention mechanism, effective noise reduction processing of the speech signal to be enhanced is realized, the call quality of the enhanced speech signal is improved, and the communication experience is enhanced. Moreover, before model training, a frequency band compression operation is added to reduce the computational complexity of the model and speed up the model training speed.
[0169] It can be understood that in the embodiments of the present application, for content related to user information, such as the content of the voice signal to be enhanced, the voice signal processing result, etc., if it involves data related to user information or enterprise information, when the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, or these information need to be blurred to eliminate the corresponding relationship between the information and the user; and the collection and processing of relevant data should strictly comply with the requirements of relevant national laws and regulations during practical applications, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data usage and processing behaviors within the scope authorized by laws and regulations and the personal information subject.
[0170] Next, the exemplary structure of the voice signal processing device 455 provided in the embodiments of the present application implemented as software modules will be further described. In some embodiments, as Figure 2 shown, the software modules stored in the voice signal processing device 455 in the memory 440 may include: a feature extraction module 4551, configured to extract features from the voice signal to be enhanced to obtain the voice features to be enhanced; a feature transformation module 4552, configured to perform multiple feature transformations on the voice features to be enhanced in a loop to obtain a voice feature vector; a feature mapping module 4553, configured to perform feature mapping on the voice feature vector in a jump connection manner to obtain a voice mapping feature; a feature fusion module 4554, configured to perform feature fusion on the voice feature vector and the voice mapping feature to obtain the gain information of the voice signal to be enhanced.
[0171] In some embodiments, the feature transformation module 4552 is further configured to: perform multiple first feature transformations on the voice features to be enhanced in a loop through a first feature transformation module with a specific number of channels to obtain a first voice feature vector; wherein, during the multiple first feature transformations, the specific number of channels of the first feature transformation module increases sequentially; perform at least one attention feature transformation on the first voice feature vector through an attention module based on the attention mechanism to obtain a second voice feature vector; perform multiple second feature transformations on the second voice feature vector in a loop through a second feature transformation module with a specific number of channels to obtain the voice feature vector; wherein, during the multiple second feature transformations, the specific number of channels of the second feature transformation module decreases sequentially.
[0172] In some embodiments, the first feature transformation module includes a first convolutional layer, a first normalization layer, a first activation layer, and a first attention mechanism layer; the feature transformation module 4552 is further configured to: during each first feature transformation, perform convolutional processing on the speech feature to be enhanced through the first convolutional layer to obtain a first convolutional feature; perform normalization processing on the first convolutional feature through the first normalization layer to obtain a first normalized feature; perform activation processing on the first normalized feature through the first activation layer to obtain a first activated feature; perform feature allocation on the first activated feature through the first attention mechanism layer to obtain the first speech feature vector.
[0173] In some embodiments, the second feature transformation module includes a second convolutional layer, a second normalization layer, a second activation layer, and a second attention mechanism layer; the specific number of channels of the first feature transformation module during the last first feature transformation of the speech feature to be enhanced is the same as the specific number of channels of the second feature transformation module during the first second feature transformation of the second speech feature vector, and moreover, the dimension of the speech feature vector obtained after the last second feature transformation of the second speech feature vector is the same as the dimension of the speech feature to be enhanced.
[0174] In some embodiments, the feature mapping module 4553 is further configured to: perform feature mapping on the speech feature vector in a skip connection manner through a multi-layer gated recurrent unit with a skip connection relationship to obtain a speech mapping feature; wherein, the multi-layer gated recurrent unit with a skip connection relationship has N layers; the input feature of the first-layer gated recurrent unit is the speech feature vector; the input feature of the Nth-layer gated recurrent unit is the output features of all the gated recurrent units connected before the Nth-layer gated recurrent unit and the speech feature vector; N is a positive integer greater than 1.
[0175] In some embodiments, for the Nth-layer gated recurrent unit, the feature mapping module 4553 is further configured to: obtain the output features of all the gated recurrent units connected before the Nth-layer gated recurrent unit and the speech feature vector; perform a vector summation operation on the output features of all the gated recurrent units connected before the Nth-layer gated recurrent unit and the speech feature vector to obtain the input feature of the Nth-layer gated recurrent unit; perform feature mapping on the input feature through the Nth-layer gated recurrent unit to obtain the output feature of the Nth-layer gated recurrent unit.
[0176] In some embodiments, the device further includes a signal enhancement module, and the signal enhancement module is configured to: decompress the gain information in frequency bands to obtain the decompressed gain information; perform a multiplication operation on the decompressed gain information and the power spectrum of the speech audio feature to obtain an enhanced audio feature; perform an inverse frequency-domain transformation on the enhanced audio feature to obtain an enhanced speech signal of the speech feature to be enhanced.
[0177] In some embodiments, the decompressed gain information includes a plurality of gain values; the signal enhancement module is further configured to: obtain the power value of the speech audio feature at each frequency component, and there is a preset mapping relationship between the power value at each frequency component and one gain value in the decompressed gain information; based on the preset mapping relationship, perform a multiplication operation on each gain value and the power value to obtain the enhanced audio feature.
[0178] In some embodiments, the feature extraction module 4551 is further configured to: perform a frequency-domain transformation on the speech signal to be enhanced to obtain a speech audio feature; perform audio feature extraction on the speech audio feature to obtain an audio feature vector; perform frequency-band compression on the audio feature vector to obtain the speech feature to be enhanced.
[0179] In some embodiments, the speech signal processing method is implemented by a speech signal processing model; the device further includes a model training module, and the model training module is configured to obtain sample data; the sample data includes a first type of speech sample, a second type of speech sample corresponding to the first type of speech sample, and a true gain between the first type of speech sample and the second type of speech sample; the first type of speech sample is a clean speech sample without noise, and the second type of speech sample is a speech sample obtained by adding noise to the first type of speech sample; perform feature preprocessing on the first type of speech sample and the second type of speech sample to obtain a sample speech feature to be enhanced, and input the sample speech feature to be enhanced into the speech signal processing model; through the feature transformation network of the speech signal processing model, perform multiple feature transformations on the sample speech feature to be enhanced in a loop to obtain a sample speech feature vector; through the feature mapping network of the speech signal processing model, perform feature mapping on the sample speech feature vector in a skip connection manner to obtain a sample speech mapping feature; through the feature fusion network of the speech signal processing model, perform feature fusion on the sample speech feature vector and the sample speech mapping feature to obtain a sample estimated gain corresponding to the sample speech feature to be enhanced; input the sample estimated gain and the true gain into a loss model for loss calculation to obtain a loss result; update the model parameters in the speech signal processing model based on the loss result to obtain a trained speech signal processing model.
[0180] In some embodiments, the model training module is further configured to: extract audio features from the first type of voice samples and the second type of voice samples respectively, and correspondingly obtain a first type of sample audio feature vector and a second type of sample audio feature vector; splice the first type of sample audio feature vector and the second type of sample audio feature vector to obtain a spliced sample voice feature; and compress the frequency band of the spliced sample voice feature to obtain the sample voice feature to be enhanced.
[0181] It should be noted that the description of the device in the embodiments of the present application is similar to the description of the above method embodiments, and has beneficial effects similar to those of the method embodiments, so it will not be repeated here. For the technical details not disclosed in the embodiments of this device, please refer to the description of the method embodiments of the present application for understanding.
[0182] The embodiments of the present application provide a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor will be caused to execute the voice signal processing method provided by the embodiments of the present application. For example, as Figure 3 shown in the voice signal processing method.
[0183] The embodiments of the present application provide a computer program product, which includes computer-executable instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the voice signal processing method described above in the embodiments of the present application.
[0184] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.
[0185] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, and may be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0186] By way of example, the computer-executable instructions may or may not correspond to a file in a file system and may be stored as part of a file that holds other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code).
[0187] By way of example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0188] As described above, the above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are all included in the protection scope of the present application.
Claims
1. A method for processing a voice signal, characterized in that The method includes: Performing feature extraction on the speech signal to be enhanced to obtain the speech features to be enhanced; Performing multiple feature transformations on the speech features to be enhanced in a loop to obtain a speech feature vector; Performing feature mapping on the speech feature vector by using a skip connection method to obtain a speech mapping feature; Performing feature fusion on the speech feature vector and the speech mapping feature to obtain the gain information of the speech signal to be enhanced.
2. The method according to claim 1, characterized in that, The performing multiple feature transformations on the speech features to be enhanced in a loop to obtain a speech feature vector includes: Performing multiple first feature transformations on the speech features to be enhanced in a loop through a first feature transformation module with a specific number of channels to obtain a first speech feature vector; wherein, during the multiple first feature transformations, the specific number of channels of the first feature transformation module increases successively; Performing at least one attention feature transformation on the first speech feature vector through an attention module based on an attention mechanism to obtain a second speech feature vector; Performing multiple second feature transformations on the second speech feature vector in a loop through a second feature transformation module with a specific number of channels to obtain the speech feature vector; wherein, during the multiple second feature transformations, the specific number of channels of the second feature transformation module decreases successively.
3. The method according to claim 2, wherein The first feature transformation module includes a first convolutional layer, a first normalization layer, a first activation layer, and a first attention mechanism layer; The performing multiple first feature transformations on the speech features to be enhanced in a loop through a first feature transformation module with a specific number of channels to obtain a first speech feature vector includes: During each first feature transformation, performing convolutional processing on the speech features to be enhanced through the first convolutional layer to obtain a first convolutional feature; Performing normalization processing on the first convolutional feature through the first normalization layer to obtain a first normalized feature; Performing activation processing on the first normalized feature through the first activation layer to obtain a first activated feature; Performing feature allocation on the first activated feature through the first attention mechanism layer to obtain the first speech feature vector.
4. The method according to claim 2, wherein The second feature transformation module includes a second convolutional layer, a second normalization layer, a second activation layer, and a second attention mechanism layer; The specific number of channels of the first feature transformation module during the last first feature transformation on the speech features to be enhanced is the same as the specific number of channels of the second feature transformation module during the first second feature transformation on the second speech feature vector, and moreover, the dimension of the speech feature vector obtained after the last second feature transformation on the second speech feature vector is the same as the dimension of the speech features to be enhanced.
5. The method according to claim 1, wherein The performing feature mapping on the speech feature vector by using a skip connection method to obtain a speech mapping feature includes: Performing feature mapping on the speech feature vector by using a skip connection method through a multi-layer gated recurrent unit with a skip connection relationship to obtain a speech mapping feature; Among them, the multi-layer gated recurrent unit with skip connection relationships has N layers; the input feature of the first-layer gated recurrent unit is the speech feature vector; the input feature of the Nth-layer gated recurrent unit is the output features of all the gated recurrent units connected before the Nth-layer gated recurrent unit and the speech feature vector; N is a positive integer greater than 1.
6. The method according to claim 5, characterized in that, For the Nth-layer gated recurrent unit, the feature mapping of the speech feature vector by the multi-layer gated recurrent unit with skip connection relationships in a skip connection manner includes: Obtaining the output features of all the gated recurrent units connected before the Nth-layer gated recurrent unit and the speech feature vector; Performing a vector summation operation on the output features of all the gated recurrent units connected before the Nth-layer gated recurrent unit and the speech feature vector to obtain the input feature of the Nth-layer gated recurrent unit; Performing feature mapping on the input feature through the Nth-layer gated recurrent unit to obtain the output feature of the Nth-layer gated recurrent unit.
7. The method according to claim 1, wherein The method further includes: Performing frequency band decompression on the gain information to obtain the decompressed gain information; Performing a multiplication operation on the decompressed gain information and the power spectrum of the speech audio feature to obtain an enhanced audio feature; Performing an inverse frequency domain transformation on the enhanced audio feature to obtain the enhanced speech signal of the speech feature to be enhanced.
8. The method according to claim 7, wherein The decompressed gain information includes a plurality of gain values; the performing a multiplication operation on the decompressed gain information and the power spectrum of the speech audio feature to obtain an enhanced audio feature includes: Obtaining the power value of the speech audio feature at each frequency component, and there is a preset mapping relationship between the power value at each frequency component and a gain value in the decompressed gain information; Based on the preset mapping relationship, performing a multiplication operation on each gain value and the power value to obtain the enhanced audio feature.
9. The method according to any one of claims 1 to 8, characterized in that, The obtaining the speech feature to be enhanced by performing feature extraction on the speech signal to be enhanced includes: Performing a frequency domain transformation on the speech signal to be enhanced to obtain a speech audio feature; Performing audio feature extraction on the speech audio feature to obtain an audio feature vector; Performing frequency band compression on the audio feature vector to obtain the speech feature to be enhanced.
10. The method according to any one of claims 1 to 8, characterized in that, The speech signal processing method is implemented through a speech signal processing model; the speech signal processing model is trained through the following steps: Obtaining sample data; the sample data includes a first type of speech sample, a second type of speech sample corresponding to the first type of speech sample, and the true gain between the first type of speech sample and the second type of speech sample; the first type of speech sample is a clean speech sample without noise, and the second type of speech sample is a speech sample obtained by adding noise to the first type of speech sample; Performing feature preprocessing on the first type of speech sample and the second type of speech sample to obtain a sample speech feature to be enhanced, and inputting the sample speech feature to be enhanced into the speech signal processing model; Through the feature transformation network of the speech signal processing model, perform multiple feature transformations on the sample speech features to be enhanced in a loop to obtain sample speech feature vectors; Through the feature mapping network of the speech signal processing model, perform feature mapping on the sample speech feature vectors in a skip connection manner to obtain sample speech mapped features; Through the feature fusion network of the speech signal processing model, perform feature fusion on the sample speech feature vectors and the sample speech mapped features to obtain the sample estimated gain of the sample speech features to be enhanced; Input the sample estimated gain and the true gain into the loss model for loss calculation to obtain a loss result; Update the model parameters in the speech signal processing model based on the loss result to obtain a trained speech signal processing model.
11. The method according to claim 10, characterized in that, The performing feature preprocessing on the first type of speech samples and the second type of speech samples to obtain sample speech features to be enhanced includes: Perform audio feature extraction on the first type of speech samples and the second type of speech samples respectively to obtain a first type of sample audio feature vector and a second type of sample audio feature vector; Perform feature concatenation on the first type of sample audio feature vector and the second type of sample audio feature vector to obtain sample concatenated speech features; Perform band compression on the sample concatenated speech features to obtain the sample speech features to be enhanced.
12. A voice signal processing device, characterized in that, The device includes: A feature extraction module, configured to perform feature extraction on the speech signal to be enhanced to obtain speech features to be enhanced; A feature transformation module, configured to perform multiple feature transformations on the speech features to be enhanced in a loop to obtain speech feature vectors; A feature mapping module, configured to perform feature mapping on the speech feature vectors in a skip connection manner to obtain speech mapped features; A feature fusion module, configured to perform feature fusion on the speech feature vectors and the speech mapped features to obtain the gain information of the speech signal to be enhanced.
13. An electronic device, characterized in that, Including: A memory, configured to store computer-executable instructions; A processor, configured to implement the speech signal processing method according to any one of claims 1 to 11 when executing the computer-executable instructions stored in the memory.
14. A computer-readable storage medium, characterized in that, Stored with computer-executable instructions, which when executed by a processor, implement the speech signal processing method according to any one of claims 1 to 11.
Citation Information
Cited By
Signal detection method and device, equipment, storage medium and program product
CN121071317A